API reference

Speech

Synthesise a reply, and mint a token so a browser can transcribe directly.

Both endpoints require the assistant to have voice enabled. Check channels.voice on the assistant before showing a microphone.

Synthesise speech#

HTTP
POST /v1/assistants/{assistant_id}/speech

Turns one utterance into audio, using the voice configured on the assistant and your own provider credentials.

Body#

session_idstringRequired. The conversation this belongs to. Characters are metered onto it.
textstringRequired. 1–1000 characters.
previous_textstringWhat was said immediately before. Up to 1000 characters.
next_textstringWhat comes next. Up to 1000 characters.

previous_text and next_text are not spoken. They give the synthesiser context so intonation carries across a sentence boundary — without them, each clause is read as though it stood alone, and a paragraph delivered a sentence at a time sounds like a list.

Why 1000 characters#

A sentence-at-a-time client never needs more, and it should not want more — synthesising a whole reply in one call means seconds of silence followed by a monologue.

A larger request is either a bug or an attempt to use this as a bulk synthesis API on someone else's provider account.

Request#

Terminal
curl -X POST \
  "https://studio.zeevaa.ai/api/v1/assistants/9f2b1c84-6e3a-4d17-b0c5-2e7a8f41d9b3/speech" \
  -H "Authorization: Bearer $ZEEVAA_API_KEY" \
  -H "Content-Type: application/json" \
  --output clause.mp3 \
  -d '{
    "session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48",
    "text": "We have three that match your budget.",
    "next_text": "The closest is the Marina Villa."
  }'

Response#

Audio bytes.

HTTP
HTTP/1.1 200 OK
Content-Type: audio/mpeg
X-Zeevaa-Characters: 37
X-Request-Id: req_01HB

X-Zeevaa-Characters is what was billed for this call. It also lands on the session and on your key's usage.

Errors#

409 voice_disabledVoice is switched off on this assistant.
409 session_endedThe session is closed or expired.
422 invalid_requestText missing, empty, or over 1000 characters.
429 rate_limitedSpeech has its own tighter ceiling. See Rate limits.
503 voice_unavailablePublished with voice on but no voice credential, or the provider rejected the request. Check Credentials in Studio.

Mint a transcription token#

HTTP
POST /v1/assistants/{assistant_id}/transcription-token

Returns a short-lived token the browser uses to open a transcription socket directly with the provider. Microphone audio never passes through your server or ours.

Body#

session_idstringRequired. The conversation this belongs to.

Request#

Terminal
curl -X POST \
  "https://studio.zeevaa.ai/api/v1/assistants/9f2b1c84-6e3a-4d17-b0c5-2e7a8f41d9b3/transcription-token" \
  -H "Authorization: Bearer $ZEEVAA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"session_id": "c31f9a70-84b2-4e05-9d6c-1a7f3b2e6d48"}'

Response#

JSON
{
  "token": "…",
  "provider": "elevenlabs",
  "model": "scribe_v1",
  "url": "wss://api.elevenlabs.io/v1/speech-to-text/realtime",
  "sample_rate": 16000,
  "expires_at": "2026-08-04T09:29:00.000Z"
}

url and sample_rate are what to connect to and what audio format to send.

Important

The token is single-use. It is consumed the moment a socket opens with it, so mint a fresh one for every connection — not one per conversation. If a visitor's connection drops and reconnects, that is a second token.

expires_at bounds an unused token. Re-mint rather than holding one: a token minted at the start of a long call will have expired by the time a reconnect needs it.

Your long-lived provider key never leaves our server; this token does, and it is single-use and short-lived precisely because it does.

Errors#

409 voice_disabledVoice is switched off on this assistant.
409 session_endedThe session is closed or expired.
429 rate_limitedSee Rate limits.
503 voice_unavailableNo voice credential configured.

Putting it together#

TypeScript
// 1. once per socket — the token is consumed on connect
const { token, url } = await zeevaa("/assistants/{id}/transcription-token", {
  session_id: sessionId,
});

// 2. transcribe in the browser, against the provider, using `token`
const transcript = await listen(token);

// 3. ordinary chat turn, marked as spoken
const reply = await zeevaa("/assistants/{id}/chat", {
  session_id: sessionId,
  message: transcript,
  mode: "voice",
});

// 4. speak it back, one sentence at a time
for (const sentence of sentences(reply)) {
  const audio = await zeevaa("/assistants/{id}/speech", {
    session_id: sessionId,
    text: sentence,
  });
  await play(audio);
}

The full pattern, with echo cancellation and interruption handling, is under Voice.

Warning

Both endpoints spend your provider credits on every call. Whatever proxy route you put in front of them must have an authentication check of your own — otherwise it is a free speech API billed to you. See securing the proxy.

Next#

  • Voice — the whole loop, and the parts that are easy to get wrong
  • Usage — what speech cost you