Core concepts

Voice

Speaking and listening — how a spoken conversation is assembled from the same session as a typed one.

Voice is not a separate product with a separate transcript. It is the same assistant, the same session and the same turn loop, reached by speaking instead of typing.

A person can say one thing and type the next. The context is unbroken, because there is only ever one conversation.

The shape of a spoken turn#

Code
1.  microphone  →  transcription  →  text
2.  text        →  POST /chat     →  streamed reply
3.  reply       →  POST /speech   →  audio  →  speaker

Step 2 is the ordinary chat endpoint with "mode": "voice". Steps 1 and 3 are the two speech endpoints.

Speaking: one sentence at a time#

The instinct is to wait for the whole reply, then synthesise it. Don't.

A four-sentence answer takes several seconds to generate. Synthesising it in one call means several seconds of silence, then a paragraph — which to a listener is indistinguishable from a broken connection.

Instead, cut the stream at sentence boundaries and synthesise each clause as it completes. The assistant starts speaking after the first sentence, while the rest is still being written.

TypeScript
let buffer = "";

for await (const event of chatStream) {
  if (event.type !== "message.delta") continue;

  buffer += event.text;

  // A sentence ends at . ! or ? followed by a space — good enough, and cheap.
  const match = buffer.match(/^(.+?[.!?])\s+(.*)$/s);
  if (!match) continue;

  const [, sentence, rest] = match;
  buffer = rest;

  await speak(sentence);
}

if (buffer.trim()) await speak(buffer);

POST /v1/assistants/{assistant_id}/speech accepts up to 1000 characters and returns audio.

JSON
{
  "session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
  "text": "We have three that match your budget.",
  "previous_text": "Let me check.",
  "next_text": "The closest is the Marina Villa."
}

previous_text and next_text are optional and worth sending. They give the synthesiser the surrounding context, so intonation carries across a sentence boundary instead of each clause being read as though it stood alone.

The response is audio/mpeg. Play it, queue the next one, don't overlap them.

Listening#

Transcription runs in the browser, against the provider directly. Your server does not proxy microphone audio, and neither do we.

Terminal
POST /v1/assistants/{assistant_id}/transcription-token
JSON
{
  "token": "…",
  "provider": "elevenlabs",
  "url": "wss://api.elevenlabs.io/v1/speech-to-text/realtime",
  "sample_rate": 16000,
  "expires_at": "2026-08-04T09:29:00.000Z"
}

Hand that token to the page and open the socket from there.

One token per socket. It is consumed on connect, so a reconnect needs a fresh one — do not mint at the start of a call and hold it.

The reason is latency. Audio is bulky and continuous, and every hop it makes adds delay to a path where delay is the whole quality metric. A short-lived token costs one request; proxying the stream would cost every millisecond of the conversation.

Your long-lived provider key never leaves our server, and the minted token expires in minutes.

Echo cancellation is required#

If the microphone is open while the assistant is speaking, it hears the assistant. Without cancellation it transcribes its own output and answers itself.

TypeScript
const stream = await navigator.mediaDevices.getUserMedia({
  audio: {
    echoCancellation: true,
    noiseSuppression: true,
    autoGainControl: true,
  },
});

Every current browser supports this. It is not optional — a voice interface without it does not work, and the failure is bizarre rather than obvious.

Interruptions#

People interrupt. If the microphone stays open while the assistant talks, a listener can cut in, and the interface should let them:

  1. Stop playback immediately
  2. Drop whatever audio is queued
  3. Abort the in-flight chat request

The partial reply is still stored, so the transcript records what was actually said before the interruption rather than what would have been said.

Marking a turn as spoken#

Send "mode": "voice" on the chat request when the message came from speech.

JSON
{
  "session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
  "message": "what's available under two million",
  "mode": "voice"
}

Two things change.

The transcript records it. Each turn stores whether it was spoken or typed, so a mixed conversation reads back accurately.

The assistant answers differently. Spoken replies avoid formatting that cannot be heard — no bullet lists, no tables, no markdown — and stay shorter, because a paragraph that scans well on screen is a monologue out loud.

Cost, and why it is limited#

Speech spends your provider credits on every call. The endpoints are therefore the most tightly limited on the platform, and every character synthesised is metered onto the session and onto your key.

Check API → Usage if a bill surprises you; it breaks down by key and by assistant.

Warning

Your proxy route sits in front of these endpoints. If it is reachable without a check of your own, anyone who finds it has a free text-to-speech API billed to you. Put a session check on it. See Next.js integration.

Latency, honestly#

Because your key stays server-side, speech requests take one extra hop:

Code
browser  →  your server  →  Zeevaa  →  provider

Expect roughly 80–150ms more before the first spoken clause than a direct call would take. Sentence-level chunking is what keeps that from being noticeable — the delay lands once, at the start, rather than on every sentence.

If voice is switched off#

channels.voice on the assistant tells you whether voice is enabled. If it is false, both speech endpoints return 409 — so read it and hide the microphone rather than showing a button that fails.

Next#