API reference

Chat

Send one message, get one reply — streamed, or as a single JSON body.

HTTP
POST /v1/assistants/{assistant_id}/chat

One turn of a conversation. The assistant may search knowledge, filter your catalogue, call your webhooks, or simply answer.

Body#

session_idstringRequired. From create session.
messagestringRequired. What the person said. 1–8000 characters.
modestringtext or voice. Defaults to text. See Voice.
streambooleanDefaults to true.

That is the whole request. There is no messages array — the session holds the transcript, and that is deliberate.

Headers#

X-Zeevaa-Format: sseNamed server-sent events. Use this outside JavaScript.
(omitted)The AI SDK stream protocol, for useChat.

Ignored when stream is false.

Streaming#

Terminal
curl -N -X POST \
  "https://studio.zeevaa.ai/api/v1/assistants/9f2b1c84-6e3a-4d17-b0c5-2e7a8f41d9b3/chat" \
  -H "Authorization: Bearer $ZEEVAA_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Zeevaa-Format: sse" \
  -d '{
    "session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
    "message": "What do you have with two bedrooms under 2M?"
  }'
HTTP
event: message.start
data: {"message_id":"msg_01HA"}

event: tool.call
data: {"name":"search_records","input":{"bedrooms":2,"price":{"lte":2000000}}}

event: canvas
data: {"kind":"results","status":"loading","expected":6}

event: tool.result
data: {"name":"search_records","status":"success","duration_ms":38}

event: canvas
data: {"kind":"results","status":"ready","total":3,"records":[]}

event: message.delta
data: {"text":"We have three that match"}

event: message.delta
data: {"text":" — the Marina Villa is closest to your budget."}

event: done
data: {"usage":{"input_tokens":412,"output_tokens":88,"total_tokens":500},"remaining_seconds":548,"finish_reason":"stop"}

Every event is documented under Streaming.

Non-streaming#

Send "stream": false for one JSON body when the turn finishes.

Terminal
curl -X POST \
  "https://studio.zeevaa.ai/api/v1/assistants/9f2b1c84-6e3a-4d17-b0c5-2e7a8f41d9b3/chat" \
  -H "Authorization: Bearer $ZEEVAA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
    "message": "And the cheapest?",
    "stream": false
  }'
JSON
{
  "message": {
    "id": "msg_01HB",
    "role": "assistant",
    "content": "The Marina Villa at 1.85M is the lowest of the three.",
    "created_at": "2026-08-04T09:15:02.771Z"
  },
  "tool_calls": [
    {
      "name": "search_records",
      "status": "success",
      "duration_ms": 31,
      "input": { "bedrooms": 2, "sort": "price_asc", "limit": 1 }
    }
  ],
  "canvas": {
    "kind": "results",
    "status": "ready",
    "total": 3,
    "records": [
      {
        "id": "8f21c07d-2b44-4e19-9a63-77c1e0d5b8aa",
        "title": "Marina Villa",
        "subtitle": "Dubai Marina · Ready 2027",
        "image_url": "https://example.com/marina.jpg",
        "data": { "price": 1850000, "bedrooms": 2 }
      }
    ]
  },
  "suggestions": ["What are the service charges?"],
  "control": null,
  "usage": { "input_tokens": 508, "output_tokens": 34, "total_tokens": 542 },
  "session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
  "remaining_seconds": 512,
  "finish_reason": "stop"
}

canvas, suggestions and control are null when the turn produced none — they carry the same payloads as the streamed events of the same names. control being non-null means the assistant has ended the conversation or asked for a human; close the session rather than sending another turn into it.

Tip

Non-streaming is the right choice for backend work — summarising a conversation, running a scheduled check, feeding a reply into another system. For anything a person is waiting on, stream. The wait is the same length; it just does not feel like it.

finish_reason#

stopThe assistant finished its answer. The ordinary case.
lengthIt hit the configured output ceiling mid-sentence.
tool_stepsIt used its maximum number of tool steps in one turn. Usually a sign the question needs narrowing.
content_filterThe provider's safety filter stopped it.
errorIt failed part-way. Whatever was generated is still stored.
abortedYou cancelled the request. The partial reply is still stored.
otherThe provider reported something none of the above covers.

One turn at a time#

Turns within a session are sequential. A second request while one is streaming returns 409 turn_in_progress — the reply would be answering a question the first turn has not finished, and both would be wrong.

Wait for the stream to close, or abort the first.

Different sessions run concurrently without limit, up to your rate limits.

Timing#

A turn is bounded at 90 seconds and almost never approaches it. Ordinary replies land in one to three seconds; a turn that searches a catalogue and then answers is typically under six.

Allow at least 120 seconds of read timeout in your client. The default in many HTTP libraries is far shorter and will cut off legitimate turns.

Errors#

404 assistant_not_foundUnknown, or not reachable by this key.
404 session_not_foundUnknown session, or it belongs to another assistant.
409 assistant_not_publishedDraft or disabled.
409 session_endedThe session is closed. Open a new one.
409 session_expiredIt ran past max_call_seconds. Open a new one.
409 turn_in_progressA turn is already streaming on this session.
422 invalid_requestMissing or malformed fields.
429 rate_limitedSee Rate limits.
503 assistant_unavailableThe assistant is misconfigured — usually a missing or deleted model credential. Fix it in Studio.

Once a stream has started, a failure arrives as an error event rather than a status code. The response was already 200 by then.

Next#

  • Speech — synthesis and transcription
  • Streaming — every event in detail