API reference

Usage

Requests, tokens and latency — by day, by key, by assistant, by endpoint.

HTTP
GET /v1/usage

Aggregated usage for the key's workspace. The same numbers behind Studio → API → Usage, so you can put them in your own dashboard, alert on them, or bill your own customers from them.

Query parameters#

fromstringISO 8601 date or timestamp. Defaults to 30 days ago.
tostringISO 8601. Defaults to now.
group_bystringday, key, assistant or endpoint. Defaults to day.
assistant_idstringRestrict to one assistant.
api_key_idstringRestrict to one key.

The window is capped at 90 days. A longer range returns 422.

Request#

Terminal
curl "https://studio.zeevaa.ai/api/v1/usage?from=2026-07-01&to=2026-08-01&group_by=day" \
  -H "Authorization: Bearer $ZEEVAA_API_KEY"

Response#

JSON
{
  "from": "2026-07-01T00:00:00.000Z",
  "to": "2026-08-01T00:00:00.000Z",
  "group_by": "day",
  "totals": {
    "requests": 18422,
    "errors": 63,
    "sessions": 3105,
    "input_tokens": 4210556,
    "output_tokens": 812900,
    "total_tokens": 5023456,
    "tts_characters": 184220,
    "stt_seconds": 9840.5,
    "p50_duration_ms": 1180,
    "p95_duration_ms": 4620
  },
  "buckets": [
    {
      "key": "2026-07-01",
      "requests": 612,
      "errors": 2,
      "sessions": 103,
      "input_tokens": 140218,
      "output_tokens": 27110,
      "total_tokens": 167328,
      "tts_characters": 6140,
      "stt_seconds": 328.4,
      "p50_duration_ms": 1140,
      "p95_duration_ms": 4380
    }
  ]
}

buckets[].key is what the grouping produced: a date for day, an id for key and assistant, a path for endpoint. When it is an id, label carries the human name beside it.

JSON
{ "key": "3f8a…", "label": "Production backend", "requests": 14208 }

What the numbers mean#

requestsEvery authenticated call, including ones that returned an error.
errorsRequests that returned 4xx or 5xx. Watch the ratio, not the count.
sessionsConversations opened in the window.
input_tokens / output_tokensModel tokens across every turn.
tts_charactersCharacters synthesised.
stt_secondsAudio transcribed.
p50_duration_ms / p95_duration_msServer-side request duration. p95 is the one to watch — p50 stays flat while the slow tail gets much worse.

This is metering, not a bill#

Model tokens are spent against your own provider credentials — the OpenAI, Anthropic, Google or ElevenLabs keys you added under Credentials. The same is true of synthesis and transcription.

So these figures tell you what your integration consumed and where. They are not an invoice from us, and they will not match a provider's dashboard exactly: providers meter with their own tokenisers and their own rounding.

Use them for attribution — which key, which assistant, which day — and use your provider's dashboard for the money.

Retention#

Detailed request logs are kept for 30 days. Aggregates outlive them, so a query for last quarter returns totals even where the underlying rows have gone.

Usage totals also outlive transcripts. A conversation ageing out of an assistant's retention window does not change last month's request count.

Errors#

422 invalid_rangefrom is after to, or the window exceeds 90 days.
422 invalid_requestUnknown group_by value.

Next#

  • Rate limits — the ceilings these numbers run into
  • Errors — what the errors count is made of