Reference

Rate limits

What each key gets, how to read the headers, and what to do at the ceiling.

Limits are per API key, not per workspace. Two keys have two budgets, which is one of the better reasons to use more than one — a runaway script on a staging key cannot throttle production.

The ceilings#

Requests60 per minute, per key
Concurrent turns10 in flight, per workspace
Speech30 per minute, per key
Transcription tokens10 per minute, per key

Speech is lower because every call spends your provider credits directly. A sentence-at-a-time voice client sits comfortably inside 30 a minute; anything approaching it is synthesising more than a conversation needs.

Need more? Ask. These are defaults, not architecture.

Headers#

Every authenticated response carries them, successful ones included — so you can back off before you are refused rather than after.

They are absent on a 401, and necessarily so: limits are counted per key, and a request we could not authenticate has no key to count against.

HTTP
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 54
X-RateLimit-Reset: 1785312840

X-RateLimit-Reset is a Unix timestamp in seconds.

When you hit one#

HTTP
HTTP/1.1 429 Too Many Requests
Retry-After: 12
JSON
{
  "error": {
    "type": "rate_limit_error",
    "code": "rate_limited",
    "message": "Too many requests. Retry in 12 seconds."
  },
  "request_id": "req_01HC"
}

Wait Retry-After seconds, then retry. Do not retry immediately, and do not retry in a tight loop — both extend the window you are locked out for.

TypeScript
if (response.status === 429) {
  const wait = Number(response.headers.get("retry-after") ?? 5);
  await new Promise((resolve) => setTimeout(resolve, wait * 1000));
  return retry();
}

Staying under them#

Queue turns per conversation. Turns in one session are sequential anyway — a second request while one is streaming returns 409, not 429. Waiting for the stream to close is both correct and free.

Don't poll transcripts. The stream tells you when a turn is done. Fetch a transcript when someone opens a history view, not on a timer.

Cache the assistant. GET /assistants/{id} changes when you change it in Studio, which is rarely. Fetch it once per page load, not once per message.

Mint transcription tokens per socket. Not per utterance. One token covers a whole connection — it is only consumed when a socket opens, so a call that stays connected needs exactly one.

Spread scheduled work. A nightly job that fires four hundred requests at exactly midnight will be throttled. Add jitter, or a small delay between them.

If the same assistant also serves a shared link or an embedded widget, that traffic is governed by the assistant's own limits — daily minutes, concurrent visitors, per-visitor caps — which you set in Studio.

The two budgets do not interact. Your integration cannot exhaust the public link, and a link that goes unexpectedly viral cannot throttle your product.

This matters more than it sounds. The public-link limits include a per-visitor hourly cap keyed on network address, and every request from your backend comes from the same handful of addresses. Applying it to API traffic would stop your integration after ten requests an hour.

Model provider limits#

These are our limits. Your model and voice providers have their own, and those are enforced against your account.

A provider rate limit surfaces as 503 assistant_unavailable rather than 429, because it is not our ceiling you hit. If you see those during a burst, check your provider dashboard — the fix is usually a tier upgrade there, not here.

Next#

  • Errors — retry rules for every status
  • Usage — see how close you are running