Core concepts
Streaming
Two response formats, every event they carry, and how to choose between them.
A reply streams as it is generated. Words arrive as the model produces them, tool calls appear when the assistant decides to use one, and catalogue results land the moment the search returns — usually well before the sentence describing them is finished.
Streaming is not decoration. A three-second reply that arrives all at once feels broken; the same three seconds with words appearing feels immediate.
Choosing a format#
| Server-sent events | AI SDK | |
|---|---|---|
| Header | X-Zeevaa-Format: sse | default |
| Consume from | any language | JavaScript / React |
| Parsing | named events, plain JSON | handled by useChat |
| Best for | Python, Go, PHP, Ruby, mobile, custom clients | a React app using @ai-sdk/react |
If you are writing a React interface, use the AI SDK format and skip the parsing entirely. Everywhere else, use SSE.
Both carry the same information; neither is a subset of the other.
If you want no stream at all, send "stream": false and get a single JSON body
when the turn finishes. See Chat.
Server-sent events#
Set the header:
X-Zeevaa-Format: sseThe response is text/event-stream. Every event has a name and a JSON payload
on a single line.
event: message.start
data: {"message_id":"msg_01H9"}
event: message.delta
data: {"text":"We have three"}
event: tool.call
data: {"name":"search_records","input":{"bedrooms":2,"price":{"lte":2000000}}}
event: tool.result
data: {"name":"search_records","status":"success","duration_ms":38}
event: canvas
data: {"kind":"results","status":"ready","total":3,"records":[]}
event: message.delta
data: {"text":" that match your budget."}
event: done
data: {"usage":{"input_tokens":412,"output_tokens":88},"remaining_seconds":548}Events#
message.start#
Always the first event, sent before generation begins.
{ "message_id": "msg_9f31c2a7b4e6", "session_id": "c31f9a70-…" }message_id is the id this reply will be stored under — the same id you will
later find on the turn in the transcript. Key your streaming message against it
and you never have to reconcile a local id with a server one.
Worth acting on rather than ignoring: on a turn that opens with a catalogue
search, the first message.delta can be several seconds away. This event is
what lets you swap a spinner for a message bubble immediately.
message.delta#
A fragment of the reply. Concatenate them in order; that is the whole of it.
Fragments are chunked by word rather than by token, so text arrives at a readable cadence instead of stuttering mid-word. It is also what lets a voice client cut clean sentences for speech.
tool.call#
The assistant decided to use a tool, with these inputs.
Show it or don't. It is what makes an activity indicator honest — "Searching properties…" rather than a generic spinner — and it is invaluable in a log when a reply is not what you expected.
Common names: search_knowledge, search_records, show_record,
compare_records, end_call, request_handoff, send_email, plus whatever
your own webhook tools are called.
tool.result#
That tool finished. status is success or error, with duration_ms.
A tool erroring is not a failed turn. The assistant is told, and usually recovers — by trying different filters, or by saying it could not find something. The stream continues.
canvas#
Structured results, ready to render.
This is the event that makes the API worth using over an embedded chat window. When the assistant searches your catalogue, you receive the actual records — not a paragraph describing them.
{
"kind": "results",
"status": "ready",
"title": "Two bedrooms under 2M",
"total": 3,
"records": [
{
"id": "8f21c07d-2b44-4e19-9a63-77c1e0d5b8aa",
"title": "Marina Villa",
"subtitle": "Dubai Marina · Ready 2027",
"image_url": "https://example.com/marina.jpg",
"data": { "price": 1850000, "bedrooms": 2, "size_sqft": 1420 }
}
]
}kind is one of:
results | A set of matching records. total is the full count; records is the page returned. |
detail | One record, expanded. |
compare | Several records, to show side by side. |
knowledge | Passages from your documents, each with a similarity score. |
empty | Nothing matched. |
status is loading, ready or error. A loading event arrives the instant
a search starts, carrying expected — how many results were asked for — so you
can paint that many skeleton cards and have the layout stay still when the real
ones land.
Note
data contains only the fields you approved for display on that catalogue.
Columns you did not tick in Studio never leave the server, whichever door the
question came through.
suggestions#
Tappable follow-ups the assistant thinks are worth offering next.
{ "replies": ["What are the service charges?", "Can I see the floor plan?"] }Chat only — they make no sense read aloud, so a turn sent with mode: "voice"
does not produce them. Optional to show, and one of the cheapest ways to keep a
conversation going.
control#
The assistant has decided the conversation is over, or that it should go to a human.
{
"end_call": { "reason": "handoff", "farewell": "Let me put you through." },
"handoff": { "message": "Our team will call you back.", "email": "sales@example.com" }
}end_call.reason is assistant_ended or handoff. Either way the assistant
has said its goodbye — close the session rather than leaving it open on a
conversation that is finished. handoff carries whatever contact details the
operator configured, so you can show them.
error#
Something went wrong mid-turn. Carries a message safe to show and a code.
The stream closes after it. See Errors.
done#
The turn is complete.
{
"usage": { "input_tokens": 412, "output_tokens": 88, "total_tokens": 500 },
"remaining_seconds": 548,
"session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
"message_id": "msg_01H9",
"finish_reason": "stop"
}remaining_seconds is time left on the session's limit.
finish_reason is one of stop, length, tool_steps, content_filter,
error, aborted or other — see the
table in the chat reference for what each means.
Reading it#
Any SSE client works. In Python, without a dependency:
import json, httpx
with httpx.stream(
"POST",
f"https://studio.zeevaa.ai/api/v1/assistants/{assistant_id}/chat",
headers={
"Authorization": f"Bearer {api_key}",
"X-Zeevaa-Format": "sse",
},
json={"session_id": session_id, "message": "Two bedrooms under 2M?"},
timeout=120,
) as response:
response.raise_for_status()
event = None
for line in response.iter_lines():
if line.startswith("event: "):
event = line[7:]
elif line.startswith("data: "):
payload = json.loads(line[6:])
if event == "message.delta":
print(payload["text"], end="", flush=True)
elif event == "canvas":
render_cards(payload["records"])
elif event == "done":
print(f"\n[{payload['usage']['total_tokens']} tokens]")Tip
Set a generous read timeout. A turn that searches a catalogue and then answers can take fifteen seconds before its first word, and a default five-second timeout will cut it off every time.
AI SDK format#
The default. This is the wire protocol used by
@ai-sdk/react, so a React interface needs no parsing at
all:
"use client";
import { useChat } from "@ai-sdk/react";
export function Chat({ sessionId }: { sessionId: string }) {
const { messages, sendMessage, status } = useChat({
api: "/api/zeevaa/chat",
body: { session_id: sessionId },
});
return (
<div>
{messages.map((message) => (
<p key={message.id}>
{message.parts
.filter((part) => part.type === "text")
.map((part) => part.text)
.join("")}
</p>
))}
<button onClick={() => sendMessage({ text: "Hello" })}>
{status === "streaming" ? "…" : "Send"}
</button>
</div>
);
}api points at your own proxy route, not at us — the key stays on your server.
See Next.js integration.
Catalogue results arrive as a data part with the stable id canvas, carrying
the same payload documented above. Successive writes reconcile into the same
object rather than appending, which is what makes a panel update in place —
skeletons, then cards — instead of stacking copies.
Cancellation#
Abort the request to stop a turn. The model is stopped, and whatever was generated up to that point is still stored — a cancelled turn is a short turn, not a lost one.
const controller = new AbortController();
fetch(url, { signal: controller.signal, method: "POST", body });
// the user navigated away, or pressed stop
controller.abort();Aborting does not end the session. Send another turn whenever you like.
Next#
- Voice — speech in and out
- Chat API reference — every request and response field