On this page

API reference

Build with your agent.

Your agent speaks the OpenAI-compatible chat format, so the official SDKs and most AI tools work with it as they are. Point them at our base URL, send your key, and set the model to your agent's id.

Base URL https://atelieragents.com/v1
Model your-agent-id
Key header Authorization: Bearer YOUR_API_KEY

Samples use your-agent-id and YOUR_API_KEY as placeholders. Open your portal and this page fills in your agent id.

Overview

An agent is an AI model we set up for one job of yours. Its instructions, tone and limits are built in: you send it the material to work on (an email, an invoice, a lead) and it sends back the result.

  • Base URL: https://atelieragents.com/v1
  • Model: your agent's id, for example invoice-clerk. Each key belongs to one agent, and responses always name it in model.
  • Format: JSON over HTTP, in the OpenAI chat completions shape.
  • Billing: per token, read and written. See Tokens and billing.
Endpoints
EndpointWhat it does
POST /v1/chat/completionsOpenAI-compatible chat, with optional streaming. Details
POST /v1/runSend text, get text. Made for no-code tools and short scripts. Details
GET /v1/modelsThe agent your key can call. Details
GET /v1/models/{id}One model card.
GET /v1/usageBalance, plan, included tokens and month-to-date usage. Details

Quickstart

  1. Find your key. It is in the email we sent when your agent went live, and it starts with ak_.
  2. Paste it in place of YOUR_API_KEY below.
  3. Run it. That is the whole setup.
curl https://atelieragents.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "your-agent-id",
    "messages": [
      {"role": "user", "content": "Paste the text your agent should work on here."}
    ]
  }'

The answer is in choices[0].message.content. Token counts are in usage, and the cost of the call is in the X-Cost response header.

Authentication

Send your key with every request, in either of these headers:

Headers (use one)
Authorization: Bearer YOUR_API_KEY
X-API-Key: YOUR_API_KEY
  • Keys start with ak_. Your portal shows only their first characters; the full key is in the email we sent you.
  • Keep keys server-side: in an environment variable, a secret store, or your tool's credentials. The API accepts calls from browsers (CORS), but a key placed in a web page or a mobile app can be read by anyone, and spent.
  • Use one key per integration so usage is easy to tell apart. Ask us for more keys.
  • We cannot show a key again. If one is lost, write to us: we revoke it and issue a new one. If one was exposed, tell us: we also reset your portal link, which signs out every open portal session, and email you the new link.
  • A missing, mistyped or revoked key gets 401 with the code invalid_api_key.

Chat completions

POST https://atelieragents.com/v1/chat/completions

The request and response follow the OpenAI chat completions format, so existing code keeps working with a new base URL. The model you send is accepted and ignored: your key decides which agent answers, and the response names it.

Request body
{
  "model": "your-agent-id",
  "messages": [
    {
      "role": "system",
      "content": "Answer in German."
    },
    {
      "role": "user",
      "content": "Invoice text goes here."
    }
  ],
  "max_tokens": 400,
  "temperature": 0.2
}
Response
{
  "id": "chatcmpl-7c1e",
  "object": "chat.completion",
  "created": 1790000000,
  "model": "your-agent-id",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Total due: 1,250.00"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 412,
    "completion_tokens": 38,
    "total_tokens": 450
  }
}
X-Tokens-Used 450X-Cost 0.0054X-Balance 42.0946X-Included-Remaining 0

Parameters

Request parameters
ParameterTypeNotes
messages requiredarrayRoles system, developer, user, assistant, tool. Content is a string or an array of content parts. At least one message that is not a system message.
modelstringOptional and ignored. Responses return your agent's id.
max_tokens, max_completion_tokensinteger, 1 or moreLongest answer, in tokens. Capped at your agent's maximum, which is also the default.
temperaturenumber, 0 to 2Defaults to your agent's setting.
top_pnumber, 0 to 1Defaults to your agent's setting.
stream, stream_optionsboolean, objectSee Streaming.
stopstring or arraySequences where the answer stops.
seedintegerPassed through.
presence_penalty, frequency_penaltynumber, -2 to 2Passed through.
response_formatobjectPassed through. See Tools and JSON output.
tools, tool_choicearray, string or objectPassed through. See Tools and JSON output.
Anything elseIgnored. n is always 1.

Request bodies are limited to 2 MB. Out-of-range values get 400 with the code invalid_parameter.

How your agent's instructions are applied

Your agent's own instructions are always sent first, and a call cannot remove or replace them. If you send system (or developer) messages, their text is added after the agent's instructions, under the heading "Additional instructions from the caller". Use them for context that changes from call to call: the language to answer in, today's date, a customer's name.

Adding context for one call
reply = client.chat.completions.create(
    model="your-agent-id",
    messages=[
        {"role": "system", "content": "Answer in German. Today is 2026-09-26."},
        {"role": "user", "content": "Paste the text your agent should work on here."},
    ],
    max_tokens=400,
    temperature=0.2,
)

The built-in instructions travel with every call, so they count as prompt tokens.

Tools and JSON output

response_format, tools and tool_choice reach your agent unchanged. Whether an answer follows a JSON format or calls one of your tools depends on how your agent was built: if your integration relies on it, tell us and we test it with you.

Asking for JSON
import json

reply = client.chat.completions.create(
    model="your-agent-id",
    messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
    response_format={"type": "json_object"},
)
data = json.loads(reply.choices[0].message.content)

Tool calls come back in choices[0].message.tool_calls with finish_reason "tool_calls". Send each result back as a message with the role tool and the matching tool_call_id.

Streaming

Set "stream": true to receive the answer while it is written, as server-sent events. Each event is a line that starts with data: followed by a JSON chunk, and the stream ends with data: [DONE].

stream = client.chat.completions.create(
    model="your-agent-id",
    messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:  # last chunk: token counts for the whole answer
        print("\nTokens used:", chunk.usage.total_tokens)
What the stream looks like
data: {"id":"chatcmpl-7c1e","object":"chat.completion.chunk","created":1790000000,"model":"your-agent-id","choices":[{"index":0,"delta":{"content":"Total"},"finish_reason":null}]}

data: {"id":"chatcmpl-7c1e","object":"chat.completion.chunk","created":1790000000,"model":"your-agent-id","choices":[{"index":0,"delta":{"content":" due: 1,250.00"},"finish_reason":null}]}

data: {"id":"chatcmpl-7c1e","object":"chat.completion.chunk","created":1790000000,"model":"your-agent-id","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"chatcmpl-7c1e","object":"chat.completion.chunk","created":1790000000,"model":"your-agent-id","choices":[],"usage":{"prompt_tokens":412,"completion_tokens":38,"total_tokens":450}}

data: [DONE]

Token counts and cost when streaming

  • Add "stream_options": {"include_usage": true} to receive a last chunk with empty choices and the usage of the whole answer. When you ask for it, it is sent at the end of every stream, except when the stream is interrupted (see below).
  • Streamed responses cannot carry the X-Tokens-Used, X-Cost and X-Balance headers: headers leave before the answer is written. Read the final usage chunk, or call GET /v1/usage for your balance.
  • If you close the connection early, the tokens generated so far are billed.
  • A stream may start with comment lines (: keepalive) while your agent prepares its answer, and they can also come between chunks. SSE clients and the OpenAI SDKs ignore them.
  • If the agent fails before the stream starts, you get a normal JSON error with an HTTP status (see Errors). If it fails after the stream has started, the stream ends with an error event and data: [DONE], without a usage chunk. Its code tells what happened: upstream_unavailable, upstream_timeout or server_busy when no token was sent yet (nothing billed), upstream_interrupted when your agent stopped answering midway, or internal_error for an unexpected error on our side. After an interruption the tokens produced so far are billed: your portal and GET /v1/usage show them.
An interrupted stream ends like this
data: {"error":{"message":"The agent is temporarily unavailable. Please retry.","type":"server_error","code":"upstream_interrupted"}}

data: [DONE]

Simple run endpoint

POST https://atelieragents.com/v1/run

The shortest path from a no-code tool or a script: send text, get text back in output, with the cost of the call in the same response.

Run request fields
FieldTypeNotes
inputstring, object or arrayThe material for your agent. Objects and arrays are sent to it as JSON text. Required unless you send messages.
messagesarrayInstead of input: the chat completions format.
max_tokens, temperatureinteger, numberOptional, same rules as chat completions.
curl https://atelieragents.com/v1/run \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"input": "Paste the text your agent should work on here."}'
Response
{
  "id": "run_5f0c2a9e1b7d4c3a8e6f1d20",
  "object": "agent.run",
  "created": 1790000000,
  "agent": "your-agent-id",
  "output": "Total due: 1,250.00",
  "finish_reason": "stop",
  "usage": {
    "prompt_tokens": 412,
    "completion_tokens": 38,
    "total_tokens": 450
  },
  "billing": {
    "cost": "0.0054",
    "currency": "EUR",
    "included_tokens_used": 0,
    "included_tokens_remaining": 0,
    "balance": "42.0946"
  }
}
  • billing.cost and billing.balance are strings in EUR, included_tokens_remaining counts the included tokens left in your period. In rare cases billing is null; the answer is still valid.
  • Streaming, tools and response_format are not available here. Use chat completions for them.

Models

GET https://atelieragents.com/v1/models and /v1/models/{id}

The list holds one model: the agent your key calls. Its id is the value to send as model. /v1/models/{id} returns that card, or 404 model_not_found for any other id.

Request
curl https://atelieragents.com/v1/models \
  -H "Authorization: Bearer YOUR_API_KEY"
Response
{
  "object": "list",
  "data": [
    {
      "id": "your-agent-id",
      "object": "model",
      "created": 1788000000,
      "owned_by": "atelier"
    }
  ]
}

The models and usage endpoints still answer when your balance is empty, your account is suspended or your agent is paused, so your tools can always check where you stand.

Usage and balance

GET https://atelieragents.com/v1/usage

Your balance, plan, included tokens and usage for the current month, in one call. It does not cost anything.

Request
curl https://atelieragents.com/v1/usage \
  -H "X-API-Key: YOUR_API_KEY"
Response (example values)
{
  "object": "usage",
  "agent": "your-agent-id",
  "plan": {
    "slug": "starter",
    "name": "Starter",
    "kind": "subscription"
  },
  "currency": "EUR",
  "balance": "42.10",
  "credit_limit": "0.00",
  "available": "42.10",
  "included_tokens_remaining": 3650000,
  "period_end": "2026-11-12T09:00:00Z",
  "rates": {
    "input_per_million": "10.00",
    "output_per_million": "10.00"
  },
  "month_to_date": {
    "since": "2026-10-01T00:00:00Z",
    "requests": 257,
    "prompt_tokens": 1042200,
    "completion_tokens": 307800,
    "total_tokens": 1350000,
    "cost": "0.00"
  },
  "portal": "https://atelieragents.com/portal",
  "generated_at": "2026-10-08T23:13:18Z"
}
Usage fields
FieldMeaning
balance, credit_limit, availableAmounts in EUR, as strings. available is your balance plus any agreed credit line.
included_tokens_remainingIncluded tokens left in the current period (0 on pay as you go).
period_endEnd of the current plan period, in UTC. null on pay as you go.
ratesYour price per 1M prompt tokens (input_per_million) and answer tokens (output_per_million).
month_to_dateCalls, tokens and cost since the 1st of the month, UTC.

Response headers

Billing headers
HeaderMeaningExample
X-Tokens-UsedPrompt plus answer tokens of this call.450
X-CostCost of this call, in EUR.0.0054
X-BalanceYour balance after this call.42.0946
X-Included-RemainingIncluded tokens left in your period.0
X-AI-GeneratedEvery answer is marked as generated by an AI, in a form a program can read (also "ai_generated": true in /v1/run).true

Sent with every billed, non-streamed answer from /v1/chat/completions and /v1/run. Browsers can read them too (they are exposed through CORS).

raw = client.chat.completions.with_raw_response.create(
    model="your-agent-id",
    messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
)
print(raw.headers.get("x-cost"), raw.headers.get("x-balance"))
reply = raw.parse()  # the usual ChatCompletion object

Usage export (CSV)

The Download CSV button on the Usage page of your portal gives one row per call for the period you picked. The English file uses commas and a decimal point, with the column names below. The French portal gives the same columns with French headers, semicolons and decimal commas, so it opens in columns in French spreadsheet apps.

Usage CSV columns
ColumnMeaning
time_utcWhen the call was made, in UTC.
agentYour agent's id, the model name you send.
endpointThe endpoint called: chat.completions for /v1/chat/completions, run for /v1/run.
streamyes when the answer was streamed.
statusHTTP status of the call: 200 when it worked, otherwise see Errors. 499 means the caller closed the stream before the end: the tokens produced so far are billed.
prompt_tokens, completion_tokens, total_tokensTokens your agent read, tokens it wrote, and both together.
included_tokensTokens covered by your plan's included allowance.
billed_tokensTokens paid from your balance.
cost, currencyWhat the call took from your balance, and the currency code.
latency_msHow long the call took, in milliseconds.
estimatedyes when the token counts were estimated.

Errors

Errors use the OpenAI error shape, so SDKs raise them with our message. The code field tells you what happened.

HTTP 402
{
  "error": {
    "message": "Your balance is used up and no included tokens remain. Top up or change plan at https://atelieragents.com/portal",
    "type": "insufficient_quota",
    "code": "insufficient_quota"
  }
}
Error codes
StatusCodeMeaningWhat to do
400nullThe body is empty, is not valid JSON, or is not a JSON object.Send a JSON object with the header Content-Type: application/json.
400invalid_messages
invalid_parameter
invalid_input
The request is malformed: no message, a message in the wrong shape, a value out of range.Fix the request. Retrying it unchanged fails again.
400request_rejectedYour agent could not process the input, usually because it is too long. Not billed.Send less text, or lower max_tokens.
400unsupported_requestYour request uses something this agent cannot handle, such as a content part type or a tool definition it does not support. The message names the part. Not billed.Remove or change that part of the request.
401invalid_api_keyThe key is missing, mistyped or revoked.Check the header. Ask us for a new key if needed.
402insufficient_quotaYour balance is used up and no included tokens remain, or calls still running already use what is left.Top up or change plan in your portal. If other calls were running, retry when they finish; otherwise retrying does not help until you top up.
403account_suspended
agent_paused
Your account is suspended, or this agent is paused.Contact us.
404model_not_foundOn /v1/models/{id}: this id is not your agent. Unknown paths also get 404.Use the id from /v1/models.
413request_too_largeThe body is over 2 MB.Send less per call: split long documents.
429rate_limit_exceededToo many calls in the last minute for this key.Wait the number of seconds in Retry-After, then retry.
500nullAn unexpected error on our side.Retry with backoff.
502upstream_unavailableYour agent is temporarily unavailable. Not billed.Retry with backoff.
503server_busyYour agent is busy right now (too many calls at once). Not billed.Wait the number of seconds in Retry-After (1 to 60), then retry.
503agent_backend_unavailableYour agent's own AI server, and any fallback set for it, are unavailable right now (switched off or outside their opening hours). The request was not sent anywhere else. Not billed.Treat the item as to be checked by hand, or retry later. When one of them opens at set hours, Retry-After gives the seconds until it opens.
504upstream_timeoutYour agent took too long to answer. Not billed.Retry, or lower max_tokens and send less text.

Checks run in this order: body size, key, account and agent status, balance, rate limit, whether your agent's AI server is available, then the content of the request. Before a 502, 503 or 504, we may already have retried your call once on our side, so wait a moment before you retry.

Rate limits

Your keys have no limit on calls per minute unless we agreed on one with you, so a large batch or a busy agent is never slowed down by a counter. Your balance still applies, and your portal shows the limit, if any, next to each key. When a key has a limit, the window slides and counts the calls of the last 60 seconds.

  • Over the limit of a key that has one, a call gets 429 with a Retry-After header, in seconds.
  • /v1/models and /v1/usage have their own counter with the same limit, so checking your balance never uses up your call budget.
  • At busy moments, a call can wait a few seconds before your agent starts on it. If no slot frees up in time, it gets 503 server_busy with a Retry-After header, and nothing is billed.

Tokens and billing

You pay for tokens: pieces of words, about 4 characters of English text each. Every call counts two kinds:

  • Prompt tokens: everything your agent reads, meaning your messages plus its built-in instructions.
  • Answer tokens: everything it writes.

Your rates depend on your plan: see Pricing, or your portal for your own figures.

This covers the agents we run for you. An agent you buy once is delivered to you and runs on your side: its running costs are those of your own AI account or server.

Example

Prompt tokens
1,500
Answer tokens
500
Rate, pay as you go
€12.00 / 1M
Cost of the call
€0.024
Calculation
2,000 x €12.00 / 1,000,000

How a call is billed

  1. Included tokens first. On a plan, prompt tokens and then answer tokens come out of your monthly allowance. Unused included tokens expire at the end of each period.
  2. Then your balance. Tokens beyond the allowance, or every token on pay as you go, are paid from your prepaid balance at your plan's rates per 1M tokens.
  3. Rounded up per call to the next millionth of the currency unit (€0.000001). Your totals are the exact sum of your calls.

When the balance runs out

Calls are accepted while you have included tokens left, or while your balance plus any agreed credit line is above zero. The last call can take the balance slightly below zero; the next one gets 402 insufficient_quota.

Calls that are still running count against your balance, so when it is almost used up, parallel calls may get 402 insufficient_quota before the first one finishes. Retry when they finish.

On a plan paid online, calls keep going through for up to 7 days after the period ends while the renewal is being paid. Once it is paid, they count against the new period.

When the money you can spend (your balance plus any credit line) drops under €2.00, we email you once. The alert re-arms after you top up.

Estimated counts

If your agent does not report exact token counts for a call, we estimate them from the length of the text, at about 4 characters per token. Those calls are marked "estimated" in your portal and in the usage CSV.

SDK examples

Any OpenAI-compatible client works: set its base URL to https://atelieragents.com/v1 and the model to your agent's id.

from openai import OpenAI

client = OpenAI(
    base_url="https://atelieragents.com/v1",
    api_key="YOUR_API_KEY",
)

reply = client.chat.completions.create(
    model="your-agent-id",
    messages=[{"role": "user", "content": "Paste the text your agent should work on here."}],
)
print(reply.choices[0].message.content)
Libraries
LanguageInstallSetting
Pythonpip install openaiOpenAI(base_url=..., api_key=...)
JavaScript, TypeScriptnpm install openainew OpenAI({ baseURL, apiKey })
Anything elseAny HTTP clientPOST JSON with the key header

No-code tools

Make, Zapier, n8n and similar tools call your agent with their HTTP request step. Use the simple run endpoint with these settings:

HTTP request settings
Method        POST
URL           https://atelieragents.com/v1/run
Header        Authorization: Bearer YOUR_API_KEY
Header        Content-Type: application/json
Body (JSON)   {"input": "map the text from the previous step here"}
Result        read the "output" field of the response
Where to find the HTTP step
ToolStep to useNotes
MakeHTTP, "Make a request"Body type raw, content type JSON. Turn on response parsing, then map output into the next module.
ZapierWebhooks by Zapier, "Custom Request"Method POST, the JSON in Data, the two headers in Headers. The answer is in output.
n8nHTTP Request nodeMethod POST, send a JSON body, add the Authorization header (or a header credential). Read output.
  • Labels differ between versions of these tools; the settings above stay the same.
  • When you insert a field inside the JSON body, make sure quotes and line breaks in the text are escaped. Most tools offer a JSON-safe insert.
  • Many tools stop waiting after 30 to 60 seconds. Add "max_tokens": 400 (or what you need) to the body to keep calls short.

Best practices

  • Keep prompts short. Your agent already knows its job. Send the material to work on, not long instructions: you pay for every token it reads.
  • Cap max_tokens to the longest answer you need. It bounds both the cost and the time of each call.
  • Retry with backoff on 429, 500, 502, 503 and 504, and respect Retry-After. Do not retry 400, 401, 402 or 403.
  • Set a client timeout of one to two minutes, or stream long answers.
  • Watch your balance: read X-Balance, or call /v1/usage before a large batch.
  • One key per integration, kept out of front-end code and out of version control.
from openai import OpenAI

client = OpenAI(
    base_url="https://atelieragents.com/v1",
    api_key="YOUR_API_KEY",
    max_retries=5,  # retries 429 and 5xx with backoff and honors Retry-After
    timeout=120,    # seconds: long answers take time
)

Changelog

  1. Busy answers and French docs

    At busy moments a call waits a few seconds for a free slot; if none frees up, it gets 503 server_busy with Retry-After. This reference is now also available in French.

  2. API v1

    Chat completions with streaming and a final usage chunk, the simple run endpoint, models, usage and balance, billing headers on every non-streamed answer.

We add fields to responses over time. Ignore the fields you do not use, and your integration keeps working.