Core concepts

Streaming

Two response formats, every event they carry, and how to choose between them.

A reply streams as it is generated. Words arrive as the model produces them, tool calls appear when the assistant decides to use one, and catalogue results land the moment the search returns — usually well before the sentence describing them is finished.

Streaming is not decoration. A three-second reply that arrives all at once feels broken; the same three seconds with words appearing feels immediate.

Choosing a format#

Server-sent eventsAI SDK
HeaderX-Zeevaa-Format: ssedefault
Consume fromany languageJavaScript / React
Parsingnamed events, plain JSONhandled by useChat
Best forPython, Go, PHP, Ruby, mobile, custom clientsa React app using @ai-sdk/react

If you are writing a React interface, use the AI SDK format and skip the parsing entirely. Everywhere else, use SSE.

Both carry the same information; neither is a subset of the other.

If you want no stream at all, send "stream": false and get a single JSON body when the turn finishes. See Chat.

Server-sent events#

Set the header:

HTTP
X-Zeevaa-Format: sse

The response is text/event-stream. Every event has a name and a JSON payload on a single line.

HTTP
event: message.start
data: {"message_id":"msg_01H9"}

event: message.delta
data: {"text":"We have three"}

event: tool.call
data: {"name":"search_records","input":{"bedrooms":2,"price":{"lte":2000000}}}

event: tool.result
data: {"name":"search_records","status":"success","duration_ms":38}

event: canvas
data: {"kind":"results","status":"ready","total":3,"records":[]}

event: message.delta
data: {"text":" that match your budget."}

event: done
data: {"usage":{"input_tokens":412,"output_tokens":88},"remaining_seconds":548}

Events#

message.start#

Always the first event, sent before generation begins.

JSON
{ "message_id": "msg_9f31c2a7b4e6", "session_id": "c31f9a70-…" }

message_id is the id this reply will be stored under — the same id you will later find on the turn in the transcript. Key your streaming message against it and you never have to reconcile a local id with a server one.

Worth acting on rather than ignoring: on a turn that opens with a catalogue search, the first message.delta can be several seconds away. This event is what lets you swap a spinner for a message bubble immediately.

message.delta#

A fragment of the reply. Concatenate them in order; that is the whole of it.

Fragments are chunked by word rather than by token, so text arrives at a readable cadence instead of stuttering mid-word. It is also what lets a voice client cut clean sentences for speech.

tool.call#

The assistant decided to use a tool, with these inputs.

Show it or don't. It is what makes an activity indicator honest — "Searching properties…" rather than a generic spinner — and it is invaluable in a log when a reply is not what you expected.

Common names: search_knowledge, search_records, show_record, compare_records, end_call, request_handoff, send_email, plus whatever your own webhook tools are called.

tool.result#

That tool finished. status is success or error, with duration_ms.

A tool erroring is not a failed turn. The assistant is told, and usually recovers — by trying different filters, or by saying it could not find something. The stream continues.

canvas#

Structured results, ready to render.

This is the event that makes the API worth using over an embedded chat window. When the assistant searches your catalogue, you receive the actual records — not a paragraph describing them.

JSON
{
  "kind": "results",
  "status": "ready",
  "title": "Two bedrooms under 2M",
  "total": 3,
  "records": [
    {
      "id": "8f21c07d-2b44-4e19-9a63-77c1e0d5b8aa",
      "title": "Marina Villa",
      "subtitle": "Dubai Marina · Ready 2027",
      "image_url": "https://example.com/marina.jpg",
      "data": { "price": 1850000, "bedrooms": 2, "size_sqft": 1420 }
    }
  ]
}

kind is one of:

resultsA set of matching records. total is the full count; records is the page returned.
detailOne record, expanded.
compareSeveral records, to show side by side.
knowledgePassages from your documents, each with a similarity score.
emptyNothing matched.

status is loading, ready or error. A loading event arrives the instant a search starts, carrying expected — how many results were asked for — so you can paint that many skeleton cards and have the layout stay still when the real ones land.

Note

data contains only the fields you approved for display on that catalogue. Columns you did not tick in Studio never leave the server, whichever door the question came through.

suggestions#

Tappable follow-ups the assistant thinks are worth offering next.

JSON
{ "replies": ["What are the service charges?", "Can I see the floor plan?"] }

Chat only — they make no sense read aloud, so a turn sent with mode: "voice" does not produce them. Optional to show, and one of the cheapest ways to keep a conversation going.

control#

The assistant has decided the conversation is over, or that it should go to a human.

JSON
{
  "end_call": { "reason": "handoff", "farewell": "Let me put you through." },
  "handoff": { "message": "Our team will call you back.", "email": "sales@example.com" }
}

end_call.reason is assistant_ended or handoff. Either way the assistant has said its goodbye — close the session rather than leaving it open on a conversation that is finished. handoff carries whatever contact details the operator configured, so you can show them.

error#

Something went wrong mid-turn. Carries a message safe to show and a code. The stream closes after it. See Errors.

done#

The turn is complete.

JSON
{
  "usage": { "input_tokens": 412, "output_tokens": 88, "total_tokens": 500 },
  "remaining_seconds": 548,
  "session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
  "message_id": "msg_01H9",
  "finish_reason": "stop"
}

remaining_seconds is time left on the session's limit.

finish_reason is one of stop, length, tool_steps, content_filter, error, aborted or other — see the table in the chat reference for what each means.

Reading it#

Any SSE client works. In Python, without a dependency:

Python
import json, httpx

with httpx.stream(
    "POST",
    f"https://studio.zeevaa.ai/api/v1/assistants/{assistant_id}/chat",
    headers={
        "Authorization": f"Bearer {api_key}",
        "X-Zeevaa-Format": "sse",
    },
    json={"session_id": session_id, "message": "Two bedrooms under 2M?"},
    timeout=120,
) as response:
    response.raise_for_status()
    event = None

    for line in response.iter_lines():
        if line.startswith("event: "):
            event = line[7:]
        elif line.startswith("data: "):
            payload = json.loads(line[6:])

            if event == "message.delta":
                print(payload["text"], end="", flush=True)
            elif event == "canvas":
                render_cards(payload["records"])
            elif event == "done":
                print(f"\n[{payload['usage']['total_tokens']} tokens]")

Tip

Set a generous read timeout. A turn that searches a catalogue and then answers can take fifteen seconds before its first word, and a default five-second timeout will cut it off every time.

AI SDK format#

The default. This is the wire protocol used by @ai-sdk/react, so a React interface needs no parsing at all:

TSX
"use client";

import { useChat } from "@ai-sdk/react";

export function Chat({ sessionId }: { sessionId: string }) {
  const { messages, sendMessage, status } = useChat({
    api: "/api/zeevaa/chat",
    body: { session_id: sessionId },
  });

  return (
    <div>
      {messages.map((message) => (
        <p key={message.id}>
          {message.parts
            .filter((part) => part.type === "text")
            .map((part) => part.text)
            .join("")}
        </p>
      ))}

      <button onClick={() => sendMessage({ text: "Hello" })}>
        {status === "streaming" ? "…" : "Send"}
      </button>
    </div>
  );
}

api points at your own proxy route, not at us — the key stays on your server. See Next.js integration.

Catalogue results arrive as a data part with the stable id canvas, carrying the same payload documented above. Successive writes reconcile into the same object rather than appending, which is what makes a panel update in place — skeletons, then cards — instead of stacking copies.

Cancellation#

Abort the request to stop a turn. The model is stopped, and whatever was generated up to that point is still stored — a cancelled turn is a short turn, not a lost one.

TypeScript
const controller = new AbortController();

fetch(url, { signal: controller.signal, method: "POST", body });

// the user navigated away, or pressed stop
controller.abort();

Aborting does not end the session. Send another turn whenever you like.

Next#