API reference
Chat
Send one message, get one reply — streamed, or as a single JSON body.
POST /v1/assistants/{assistant_id}/chatOne turn of a conversation. The assistant may search knowledge, filter your catalogue, call your webhooks, or simply answer.
Body#
session_id | string | Required. From create session. |
message | string | Required. What the person said. 1–8000 characters. |
mode | string | text or voice. Defaults to text. See Voice. |
stream | boolean | Defaults to true. |
That is the whole request. There is no messages array — the session holds the
transcript, and that is deliberate.
Headers#
X-Zeevaa-Format: sse | Named server-sent events. Use this outside JavaScript. |
| (omitted) | The AI SDK stream protocol, for useChat. |
Ignored when stream is false.
Streaming#
curl -N -X POST \
"https://studio.zeevaa.ai/api/v1/assistants/9f2b1c84-6e3a-4d17-b0c5-2e7a8f41d9b3/chat" \
-H "Authorization: Bearer $ZEEVAA_API_KEY" \
-H "Content-Type: application/json" \
-H "X-Zeevaa-Format: sse" \
-d '{
"session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
"message": "What do you have with two bedrooms under 2M?"
}'event: message.start
data: {"message_id":"msg_01HA"}
event: tool.call
data: {"name":"search_records","input":{"bedrooms":2,"price":{"lte":2000000}}}
event: canvas
data: {"kind":"results","status":"loading","expected":6}
event: tool.result
data: {"name":"search_records","status":"success","duration_ms":38}
event: canvas
data: {"kind":"results","status":"ready","total":3,"records":[]}
event: message.delta
data: {"text":"We have three that match"}
event: message.delta
data: {"text":" — the Marina Villa is closest to your budget."}
event: done
data: {"usage":{"input_tokens":412,"output_tokens":88,"total_tokens":500},"remaining_seconds":548,"finish_reason":"stop"}Every event is documented under Streaming.
Non-streaming#
Send "stream": false for one JSON body when the turn finishes.
curl -X POST \
"https://studio.zeevaa.ai/api/v1/assistants/9f2b1c84-6e3a-4d17-b0c5-2e7a8f41d9b3/chat" \
-H "Authorization: Bearer $ZEEVAA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
"message": "And the cheapest?",
"stream": false
}'{
"message": {
"id": "msg_01HB",
"role": "assistant",
"content": "The Marina Villa at 1.85M is the lowest of the three.",
"created_at": "2026-08-04T09:15:02.771Z"
},
"tool_calls": [
{
"name": "search_records",
"status": "success",
"duration_ms": 31,
"input": { "bedrooms": 2, "sort": "price_asc", "limit": 1 }
}
],
"canvas": {
"kind": "results",
"status": "ready",
"total": 3,
"records": [
{
"id": "8f21c07d-2b44-4e19-9a63-77c1e0d5b8aa",
"title": "Marina Villa",
"subtitle": "Dubai Marina · Ready 2027",
"image_url": "https://example.com/marina.jpg",
"data": { "price": 1850000, "bedrooms": 2 }
}
]
},
"suggestions": ["What are the service charges?"],
"control": null,
"usage": { "input_tokens": 508, "output_tokens": 34, "total_tokens": 542 },
"session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
"remaining_seconds": 512,
"finish_reason": "stop"
}canvas, suggestions and control are null when the turn produced none —
they carry the same payloads as the
streamed events of the same names. control being
non-null means the assistant has ended the conversation or asked for a human;
close the session rather than sending another turn into it.
Tip
Non-streaming is the right choice for backend work — summarising a conversation, running a scheduled check, feeding a reply into another system. For anything a person is waiting on, stream. The wait is the same length; it just does not feel like it.
finish_reason#
stop | The assistant finished its answer. The ordinary case. |
length | It hit the configured output ceiling mid-sentence. |
tool_steps | It used its maximum number of tool steps in one turn. Usually a sign the question needs narrowing. |
content_filter | The provider's safety filter stopped it. |
error | It failed part-way. Whatever was generated is still stored. |
aborted | You cancelled the request. The partial reply is still stored. |
other | The provider reported something none of the above covers. |
One turn at a time#
Turns within a session are sequential. A second request while one is streaming
returns 409 turn_in_progress — the reply would be answering a question the
first turn has not finished, and both would be wrong.
Wait for the stream to close, or abort the first.
Different sessions run concurrently without limit, up to your rate limits.
Timing#
A turn is bounded at 90 seconds and almost never approaches it. Ordinary replies land in one to three seconds; a turn that searches a catalogue and then answers is typically under six.
Allow at least 120 seconds of read timeout in your client. The default in many HTTP libraries is far shorter and will cut off legitimate turns.
Errors#
404 assistant_not_found | Unknown, or not reachable by this key. |
404 session_not_found | Unknown session, or it belongs to another assistant. |
409 assistant_not_published | Draft or disabled. |
409 session_ended | The session is closed. Open a new one. |
409 session_expired | It ran past max_call_seconds. Open a new one. |
409 turn_in_progress | A turn is already streaming on this session. |
422 invalid_request | Missing or malformed fields. |
429 rate_limited | See Rate limits. |
503 assistant_unavailable | The assistant is misconfigured — usually a missing or deleted model credential. Fix it in Studio. |
Once a stream has started, a failure arrives as an
error event rather than a status code. The response
was already 200 by then.