On this page
API reference
Build with your agent.
Your agent speaks the OpenAI-compatible chat format, so the official SDKs and most AI tools work with it as they are. Point them at our base URL, send your key, and set the model to your agent's id.
https://atelieragents.com/v1
your-agent-id
Authorization: Bearer YOUR_API_KEY
Samples use your-agent-id and YOUR_API_KEY as placeholders. Open your portal and this page fills in your agent id.
Overview
An agent is an AI model we set up for one job of yours. Its instructions, tone and limits are built in: you send it the material to work on (an email, an invoice, a lead) and it sends back the result.
- Base URL:
https://atelieragents.com/v1 - Model: your agent's id, for example
invoice-clerk. Each key belongs to one agent, and responses always name it inmodel. - Format: JSON over HTTP, in the OpenAI chat completions shape.
- Billing: per token, read and written. See Tokens and billing.
| Endpoint | What it does |
|---|---|
POST /v1/chat/completions | OpenAI-compatible chat, with optional streaming. Details |
POST /v1/run | Send text, get text. Made for no-code tools and short scripts. Details |
GET /v1/models | The agent your key can call. Details |
GET /v1/models/{id} | One model card. |
GET /v1/usage | Balance, plan, included tokens and month-to-date usage. Details |
Quickstart
- Find your key. It is in the email we sent when your agent went live, and it starts with
ak_. - Paste it in place of
YOUR_API_KEYbelow. - Run it. That is the whole setup.
curl https://atelieragents.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "your-agent-id",
"messages": [
{"role": "user", "content": "Paste the text your agent should work on here."}
]
}'
from openai import OpenAI
client = OpenAI(
base_url="https://atelieragents.com/v1",
api_key="YOUR_API_KEY",
)
reply = client.chat.completions.create(
model="your-agent-id",
messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
)
print(reply.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://atelieragents.com/v1",
apiKey: "YOUR_API_KEY",
});
const reply = await client.chat.completions.create({
model: "your-agent-id",
messages: [{ role: "user", content: "Paste the text your agent should work on here." }],
});
console.log(reply.choices[0].message.content);
The answer is in choices[0].message.content. Token counts are in usage, and the cost of the call is in the X-Cost response header.
Authentication
Send your key with every request, in either of these headers:
Authorization: Bearer YOUR_API_KEY
X-API-Key: YOUR_API_KEY
- Keys start with
ak_. Your portal shows only their first characters; the full key is in the email we sent you. - Keep keys server-side: in an environment variable, a secret store, or your tool's credentials. The API accepts calls from browsers (CORS), but a key placed in a web page or a mobile app can be read by anyone, and spent.
- Use one key per integration so usage is easy to tell apart. Ask us for more keys.
- We cannot show a key again. If one is lost, write to us: we revoke it and issue a new one. If one was exposed, tell us: we also reset your portal link, which signs out every open portal session, and email you the new link.
- A missing, mistyped or revoked key gets
401with the codeinvalid_api_key.
Chat completions
POST https://atelieragents.com/v1/chat/completions
The request and response follow the OpenAI chat completions format, so existing code keeps working with a new base URL. The model you send is accepted and ignored: your key decides which agent answers, and the response names it.
{
"model": "your-agent-id",
"messages": [
{
"role": "system",
"content": "Answer in German."
},
{
"role": "user",
"content": "Invoice text goes here."
}
],
"max_tokens": 400,
"temperature": 0.2
}
{
"id": "chatcmpl-7c1e",
"object": "chat.completion",
"created": 1790000000,
"model": "your-agent-id",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Total due: 1,250.00"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 412,
"completion_tokens": 38,
"total_tokens": 450
}
}Parameters
| Parameter | Type | Notes |
|---|---|---|
messages required | array | Roles system, developer, user, assistant, tool. Content is a string or an array of content parts. At least one message that is not a system message. |
model | string | Optional and ignored. Responses return your agent's id. |
max_tokens, max_completion_tokens | integer, 1 or more | Longest answer, in tokens. Capped at your agent's maximum, which is also the default. |
temperature | number, 0 to 2 | Defaults to your agent's setting. |
top_p | number, 0 to 1 | Defaults to your agent's setting. |
stream, stream_options | boolean, object | See Streaming. |
stop | string or array | Sequences where the answer stops. |
seed | integer | Passed through. |
presence_penalty, frequency_penalty | number, -2 to 2 | Passed through. |
response_format | object | Passed through. See Tools and JSON output. |
tools, tool_choice | array, string or object | Passed through. See Tools and JSON output. |
| Anything else | Ignored. n is always 1. |
Request bodies are limited to 2 MB. Out-of-range values get 400 with the code invalid_parameter.
How your agent's instructions are applied
Your agent's own instructions are always sent first, and a call cannot remove or replace them. If you send system (or developer) messages, their text is added after the agent's instructions, under the heading "Additional instructions from the caller". Use them for context that changes from call to call: the language to answer in, today's date, a customer's name.
reply = client.chat.completions.create(
model="your-agent-id",
messages=[
{"role": "system", "content": "Answer in German. Today is 2026-09-26."},
{"role": "user", "content": "Paste the text your agent should work on here."},
],
max_tokens=400,
temperature=0.2,
)
The built-in instructions travel with every call, so they count as prompt tokens.
Tools and JSON output
response_format, tools and tool_choice reach your agent unchanged. Whether an answer follows a JSON format or calls one of your tools depends on how your agent was built: if your integration relies on it, tell us and we test it with you.
import json
reply = client.chat.completions.create(
model="your-agent-id",
messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
response_format={"type": "json_object"},
)
data = json.loads(reply.choices[0].message.content)
Tool calls come back in choices[0].message.tool_calls with finish_reason "tool_calls". Send each result back as a message with the role tool and the matching tool_call_id.
Streaming
Set "stream": true to receive the answer while it is written, as server-sent events. Each event is a line that starts with data: followed by a JSON chunk, and the stream ends with data: [DONE].
stream = client.chat.completions.create(
model="your-agent-id",
messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage: # last chunk: token counts for the whole answer
print("\nTokens used:", chunk.usage.total_tokens)
const stream = await client.chat.completions.create({
model: "your-agent-id",
messages: [{ role: "user", content: "Paste the text your agent should work on here." }],
stream: true,
stream_options: { include_usage: true },
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
if (chunk.usage) console.log("\nTokens used:", chunk.usage.total_tokens);
}
curl -N https://atelieragents.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "your-agent-id",
"stream": true,
"stream_options": {"include_usage": true},
"messages": [{"role": "user", "content": "Paste the text your agent should work on here."}]
}'
data: {"id":"chatcmpl-7c1e","object":"chat.completion.chunk","created":1790000000,"model":"your-agent-id","choices":[{"index":0,"delta":{"content":"Total"},"finish_reason":null}]}
data: {"id":"chatcmpl-7c1e","object":"chat.completion.chunk","created":1790000000,"model":"your-agent-id","choices":[{"index":0,"delta":{"content":" due: 1,250.00"},"finish_reason":null}]}
data: {"id":"chatcmpl-7c1e","object":"chat.completion.chunk","created":1790000000,"model":"your-agent-id","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-7c1e","object":"chat.completion.chunk","created":1790000000,"model":"your-agent-id","choices":[],"usage":{"prompt_tokens":412,"completion_tokens":38,"total_tokens":450}}
data: [DONE]
Token counts and cost when streaming
- Add
"stream_options": {"include_usage": true}to receive a last chunk with emptychoicesand theusageof the whole answer. When you ask for it, it is sent at the end of every stream, except when the stream is interrupted (see below). - Streamed responses cannot carry the
X-Tokens-Used,X-CostandX-Balanceheaders: headers leave before the answer is written. Read the final usage chunk, or callGET /v1/usagefor your balance. - If you close the connection early, the tokens generated so far are billed.
- A stream may start with comment lines (
: keepalive) while your agent prepares its answer, and they can also come between chunks. SSE clients and the OpenAI SDKs ignore them. - If the agent fails before the stream starts, you get a normal JSON error with an HTTP status (see Errors). If it fails after the stream has started, the stream ends with an error event and
data: [DONE], without a usage chunk. Itscodetells what happened:upstream_unavailable,upstream_timeoutorserver_busywhen no token was sent yet (nothing billed),upstream_interruptedwhen your agent stopped answering midway, orinternal_errorfor an unexpected error on our side. After an interruption the tokens produced so far are billed: your portal andGET /v1/usageshow them.
data: {"error":{"message":"The agent is temporarily unavailable. Please retry.","type":"server_error","code":"upstream_interrupted"}}
data: [DONE]
Simple run endpoint
POST https://atelieragents.com/v1/run
The shortest path from a no-code tool or a script: send text, get text back in output, with the cost of the call in the same response.
| Field | Type | Notes |
|---|---|---|
input | string, object or array | The material for your agent. Objects and arrays are sent to it as JSON text. Required unless you send messages. |
messages | array | Instead of input: the chat completions format. |
max_tokens, temperature | integer, number | Optional, same rules as chat completions. |
curl https://atelieragents.com/v1/run \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"input": "Paste the text your agent should work on here."}'
import requests
response = requests.post(
"https://atelieragents.com/v1/run",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={"input": "Paste the text your agent should work on here."},
timeout=120,
)
response.raise_for_status()
print(response.json()["output"])
const response = await fetch("https://atelieragents.com/v1/run", {
method: "POST",
headers: { "Authorization": "Bearer YOUR_API_KEY", "Content-Type": "application/json" },
body: JSON.stringify({ input: "Paste the text your agent should work on here." }),
});
const { output, usage, billing } = await response.json();
{
"id": "run_5f0c2a9e1b7d4c3a8e6f1d20",
"object": "agent.run",
"created": 1790000000,
"agent": "your-agent-id",
"output": "Total due: 1,250.00",
"finish_reason": "stop",
"usage": {
"prompt_tokens": 412,
"completion_tokens": 38,
"total_tokens": 450
},
"billing": {
"cost": "0.0054",
"currency": "EUR",
"included_tokens_used": 0,
"included_tokens_remaining": 0,
"balance": "42.0946"
}
}
billing.costandbilling.balanceare strings in EUR,included_tokens_remainingcounts the included tokens left in your period. In rare casesbillingisnull; the answer is still valid.- Streaming, tools and
response_formatare not available here. Use chat completions for them.
Models
GET https://atelieragents.com/v1/models and /v1/models/{id}
The list holds one model: the agent your key calls. Its id is the value to send as model. /v1/models/{id} returns that card, or 404 model_not_found for any other id.
curl https://atelieragents.com/v1/models \
-H "Authorization: Bearer YOUR_API_KEY"
{
"object": "list",
"data": [
{
"id": "your-agent-id",
"object": "model",
"created": 1788000000,
"owned_by": "atelier"
}
]
}
The models and usage endpoints still answer when your balance is empty, your account is suspended or your agent is paused, so your tools can always check where you stand.
Usage and balance
GET https://atelieragents.com/v1/usage
Your balance, plan, included tokens and usage for the current month, in one call. It does not cost anything.
curl https://atelieragents.com/v1/usage \
-H "X-API-Key: YOUR_API_KEY"
{
"object": "usage",
"agent": "your-agent-id",
"plan": {
"slug": "starter",
"name": "Starter",
"kind": "subscription"
},
"currency": "EUR",
"balance": "42.10",
"credit_limit": "0.00",
"available": "42.10",
"included_tokens_remaining": 3650000,
"period_end": "2026-11-12T09:00:00Z",
"rates": {
"input_per_million": "10.00",
"output_per_million": "10.00"
},
"month_to_date": {
"since": "2026-10-01T00:00:00Z",
"requests": 257,
"prompt_tokens": 1042200,
"completion_tokens": 307800,
"total_tokens": 1350000,
"cost": "0.00"
},
"portal": "https://atelieragents.com/portal",
"generated_at": "2026-10-08T23:15:17Z"
}
| Field | Meaning |
|---|---|
balance, credit_limit, available | Amounts in EUR, as strings. available is your balance plus any agreed credit line. |
included_tokens_remaining | Included tokens left in the current period (0 on pay as you go). |
period_end | End of the current plan period, in UTC. null on pay as you go. |
rates | Your price per 1M prompt tokens (input_per_million) and answer tokens (output_per_million). |
month_to_date | Calls, tokens and cost since the 1st of the month, UTC. |
Response headers
| Header | Meaning | Example |
|---|---|---|
X-Tokens-Used | Prompt plus answer tokens of this call. | 450 |
X-Cost | Cost of this call, in EUR. | 0.0054 |
X-Balance | Your balance after this call. | 42.0946 |
X-Included-Remaining | Included tokens left in your period. | 0 |
X-AI-Generated | Every answer is marked as generated by an AI, in a form a program can read (also "ai_generated": true in /v1/run). | true |
Sent with every billed, non-streamed answer from /v1/chat/completions and /v1/run. Browsers can read them too (they are exposed through CORS).
raw = client.chat.completions.with_raw_response.create(
model="your-agent-id",
messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
)
print(raw.headers.get("x-cost"), raw.headers.get("x-balance"))
reply = raw.parse() # the usual ChatCompletion object
const { data: reply, response } = await client.chat.completions
.create({ model: "your-agent-id", messages: [{ role: "user", content: "..." }] })
.withResponse();
console.log(response.headers.get("x-cost"), response.headers.get("x-balance"));
Usage export (CSV)
The Download CSV button on the Usage page of your portal gives one row per call for the period you picked. The English file uses commas and a decimal point, with the column names below. The French portal gives the same columns with French headers, semicolons and decimal commas, so it opens in columns in French spreadsheet apps.
| Column | Meaning |
|---|---|
time_utc | When the call was made, in UTC. |
agent | Your agent's id, the model name you send. |
endpoint | The endpoint called: chat.completions for /v1/chat/completions, run for /v1/run. |
stream | yes when the answer was streamed. |
status | HTTP status of the call: 200 when it worked, otherwise see Errors. 499 means the caller closed the stream before the end: the tokens produced so far are billed. |
prompt_tokens, completion_tokens, total_tokens | Tokens your agent read, tokens it wrote, and both together. |
included_tokens | Tokens covered by your plan's included allowance. |
billed_tokens | Tokens paid from your balance. |
cost, currency | What the call took from your balance, and the currency code. |
latency_ms | How long the call took, in milliseconds. |
estimated | yes when the token counts were estimated. |
Errors
Errors use the OpenAI error shape, so SDKs raise them with our message. The code field tells you what happened.
{
"error": {
"message": "Your balance is used up and no included tokens remain. Top up or change plan at https://atelieragents.com/portal",
"type": "insufficient_quota",
"code": "insufficient_quota"
}
}
| Status | Code | Meaning | What to do |
|---|---|---|---|
| 400 | null | The body is empty, is not valid JSON, or is not a JSON object. | Send a JSON object with the header Content-Type: application/json. |
| 400 | invalid_messagesinvalid_parameterinvalid_input | The request is malformed: no message, a message in the wrong shape, a value out of range. | Fix the request. Retrying it unchanged fails again. |
| 400 | request_rejected | Your agent could not process the input, usually because it is too long. Not billed. | Send less text, or lower max_tokens. |
| 400 | unsupported_request | Your request uses something this agent cannot handle, such as a content part type or a tool definition it does not support. The message names the part. Not billed. | Remove or change that part of the request. |
| 401 | invalid_api_key | The key is missing, mistyped or revoked. | Check the header. Ask us for a new key if needed. |
| 402 | insufficient_quota | Your balance is used up and no included tokens remain, or calls still running already use what is left. | Top up or change plan in your portal. If other calls were running, retry when they finish; otherwise retrying does not help until you top up. |
| 403 | account_suspendedagent_paused | Your account is suspended, or this agent is paused. | Contact us. |
| 404 | model_not_found | On /v1/models/{id}: this id is not your agent. Unknown paths also get 404. | Use the id from /v1/models. |
| 413 | request_too_large | The body is over 2 MB. | Send less per call: split long documents. |
| 429 | rate_limit_exceeded | Too many calls in the last minute for this key. | Wait the number of seconds in Retry-After, then retry. |
| 500 | null | An unexpected error on our side. | Retry with backoff. |
| 502 | upstream_unavailable | Your agent is temporarily unavailable. Not billed. | Retry with backoff. |
| 503 | server_busy | Your agent is busy right now (too many calls at once). Not billed. | Wait the number of seconds in Retry-After (1 to 60), then retry. |
| 503 | agent_backend_unavailable | Your agent's own AI server, and any fallback set for it, are unavailable right now (switched off or outside their opening hours). The request was not sent anywhere else. Not billed. | Treat the item as to be checked by hand, or retry later. When one of them opens at set hours, Retry-After gives the seconds until it opens. |
| 504 | upstream_timeout | Your agent took too long to answer. Not billed. | Retry, or lower max_tokens and send less text. |
Checks run in this order: body size, key, account and agent status, balance, rate limit, whether your agent's AI server is available, then the content of the request. Before a 502, 503 or 504, we may already have retried your call once on our side, so wait a moment before you retry.
Rate limits
Your keys have no limit on calls per minute unless we agreed on one with you, so a large batch or a busy agent is never slowed down by a counter. Your balance still applies, and your portal shows the limit, if any, next to each key. When a key has a limit, the window slides and counts the calls of the last 60 seconds.
- Over the limit of a key that has one, a call gets
429with aRetry-Afterheader, in seconds. /v1/modelsand/v1/usagehave their own counter with the same limit, so checking your balance never uses up your call budget.- At busy moments, a call can wait a few seconds before your agent starts on it. If no slot frees up in time, it gets
503server_busywith aRetry-Afterheader, and nothing is billed.
Tokens and billing
You pay for tokens: pieces of words, about 4 characters of English text each. Every call counts two kinds:
- Prompt tokens: everything your agent reads, meaning your messages plus its built-in instructions.
- Answer tokens: everything it writes.
Your rates depend on your plan: see Pricing, or your portal for your own figures.
This covers the agents we run for you. An agent you buy once is delivered to you and runs on your side: its running costs are those of your own AI account or server.
Example
- Prompt tokens
- 1,500
- Answer tokens
- 500
- Rate, pay as you go
- €12.00 / 1M
- Cost of the call
- €0.024
- Calculation
- 2,000 x €12.00 / 1,000,000
How a call is billed
- Included tokens first. On a plan, prompt tokens and then answer tokens come out of your monthly allowance. Unused included tokens expire at the end of each period.
- Then your balance. Tokens beyond the allowance, or every token on pay as you go, are paid from your prepaid balance at your plan's rates per 1M tokens.
- Rounded up per call to the next millionth of the currency unit (€0.000001). Your totals are the exact sum of your calls.
When the balance runs out
Calls are accepted while you have included tokens left, or while your balance plus any agreed credit line is above zero. The last call can take the balance slightly below zero; the next one gets 402 insufficient_quota.
Calls that are still running count against your balance, so when it is almost used up, parallel calls may get 402 insufficient_quota before the first one finishes. Retry when they finish.
On a plan paid online, calls keep going through for up to 7 days after the period ends while the renewal is being paid. Once it is paid, they count against the new period.
When the money you can spend (your balance plus any credit line) drops under €2.00, we email you once. The alert re-arms after you top up.
Estimated counts
If your agent does not report exact token counts for a call, we estimate them from the length of the text, at about 4 characters per token. Those calls are marked "estimated" in your portal and in the usage CSV.
SDK examples
Any OpenAI-compatible client works: set its base URL to https://atelieragents.com/ and the model to your agent's id.
from openai import OpenAI
client = OpenAI(
base_url="https://atelieragents.com/v1",
api_key="YOUR_API_KEY",
)
reply = client.chat.completions.create(
model="your-agent-id",
messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
)
print(reply.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://atelieragents.com/v1",
apiKey: "YOUR_API_KEY",
});
const reply = await client.chat.completions.create({
model: "your-agent-id",
messages: [{ role: "user", content: "Paste the text your agent should work on here." }],
});
console.log(reply.choices[0].message.content);
const response = await fetch("https://atelieragents.com/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "your-agent-id",
messages: [{ role: "user", content: "Paste the text your agent should work on here." }],
}),
});
const data = await response.json();
if (!response.ok) throw new Error(data.error.message);
console.log(data.choices[0].message.content);
curl https://atelieragents.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "your-agent-id",
"messages": [
{"role": "user", "content": "Paste the text your agent should work on here."}
]
}'
| Language | Install | Setting |
|---|---|---|
| Python | pip install openai | OpenAI(base_url=..., api_key=...) |
| JavaScript, TypeScript | npm install openai | new OpenAI({ baseURL, apiKey }) |
| Anything else | Any HTTP client | POST JSON with the key header |
No-code tools
Make, Zapier, n8n and similar tools call your agent with their HTTP request step. Use the simple run endpoint with these settings:
Method POST
URL https://atelieragents.com/v1/run
Header Authorization: Bearer YOUR_API_KEY
Header Content-Type: application/json
Body (JSON) {"input": "map the text from the previous step here"}
Result read the "output" field of the response
| Tool | Step to use | Notes |
|---|---|---|
| Make | HTTP, "Make a request" | Body type raw, content type JSON. Turn on response parsing, then map output into the next module. |
| Zapier | Webhooks by Zapier, "Custom Request" | Method POST, the JSON in Data, the two headers in Headers. The answer is in output. |
| n8n | HTTP Request node | Method POST, send a JSON body, add the Authorization header (or a header credential). Read output. |
- Labels differ between versions of these tools; the settings above stay the same.
- When you insert a field inside the JSON body, make sure quotes and line breaks in the text are escaped. Most tools offer a JSON-safe insert.
- Many tools stop waiting after 30 to 60 seconds. Add
"max_tokens": 400(or what you need) to the body to keep calls short.
Best practices
- Keep prompts short. Your agent already knows its job. Send the material to work on, not long instructions: you pay for every token it reads.
- Cap
max_tokensto the longest answer you need. It bounds both the cost and the time of each call. - Retry with backoff on 429, 500, 502, 503 and 504, and respect
Retry-After. Do not retry 400, 401, 402 or 403. - Set a client timeout of one to two minutes, or stream long answers.
- Watch your balance: read
X-Balance, or call/v1/usagebefore a large batch. - One key per integration, kept out of front-end code and out of version control.
from openai import OpenAI
client = OpenAI(
base_url="https://atelieragents.com/v1",
api_key="YOUR_API_KEY",
max_retries=5, # retries 429 and 5xx with backoff and honors Retry-After
timeout=120, # seconds: long answers take time
)
async function callAgent(body, attempts = 5) {
for (let attempt = 0; ; attempt++) {
const response = await fetch("https://atelieragents.com/v1/chat/completions", {
method: "POST",
headers: { "Authorization": "Bearer YOUR_API_KEY", "Content-Type": "application/json" },
body: JSON.stringify(body),
});
const retryable = response.status === 429 || response.status >= 500;
if (!retryable || attempt >= attempts - 1) return response;
const wait = Number(response.headers.get("retry-after")) || 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, wait * 1000));
}
}
Changelog
- Busy answers and French docs
At busy moments a call waits a few seconds for a free slot; if none frees up, it gets 503 server_busy with Retry-After. This reference is now also available in French.
- API v1
Chat completions with streaming and a final usage chunk, the simple run endpoint, models, usage and balance, billing headers on every non-streamed answer.
We add fields to responses over time. Ignore the fields you do not use, and your integration keeps working.