Review voice agent call recordings: STT word timings on Sume

After you ship a Gemini Live or other voice agent, transcribe the recordings with Sume STT: word timings, sentence segments, a 10-minute cap per request.

4 min readSume
All posts

Send each recording to Sume's STT 1.0 route as a public HTTPS audio_url and read back text plus word-level timings. Add segmentation mode sentence and you also get sentence segments, which map cleanly onto turns of a call. A request covers up to 10 minutes of audio, because duration_seconds tops out at 600, so split longer calls first.

Google's Gemini API changelog (read 2026-10-02) lists Gemini 3.8 Live GA on 2026-09-15, so more teams are shipping voice agents now and then needing to read what was said. STT is a batch job, not part of the live session.

What does the request look like?

audio_url is required. language_code is optional, for example en or ko; omit it for auto-detect. duration_seconds from 1 to 600 improves the usage reservation; omit it and Sume reserves 1 minute. Word timings are always returned, so there is no flag to ask for them.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: call-0142" \
  -d '{"audio_url": "https://media.sume.com/artifacts/example/call.wav", "duration_seconds": 240, "segmentation": {"mode": "sentence"}}'

What do I get back, and what do I not?

The completed result carries text, language fields when available, and words with start and end seconds from the audio start. The table lists what the docs I read show and what they do not.

STT 1.0 result fields. Sume OpenAPI and docs, read 2026-10-02.
ItemPresentNote
Full textYesPublic-safe text
Word timingsYesword, start, end in seconds
Sentence segmentsOn requestsegmentation.mode sentence
Speaker labelsNo field in the docs I readProvider knobs such as diarization are fixed server-side

How do I handle a 25-minute call?

Split it into pieces of 10 minutes or less with Timeline audio split, which takes a top-level url and 1 to 20 ranges and returns durable media.sume.com files. Transcribe each piece, then add each piece's start offset to its word times.

What does it cost?

The video-inspect page lists the public STT rate as $0.01 per audio minute and says to confirm it live in GET /v1/catalog.

What should I do with the transcript?

Start with the failures. Search each transcript for the phrases your agent should never say and for the points where the caller repeats themselves; the word timings let you jump to that second of the recording. Keep the job id and the original recording URL with the transcript so a reviewer can listen to the source.

Treat the text as a working record, not a verdict. Check a sample against the audio before you build metrics on it, particularly for names, numbers and mixed-language calls, and set language_code when you know the language.

How do I keep reviewers out of private audio?

Only send recordings you are allowed to process, and use a URL that is reachable by Sume when the job runs. Sume's docs prefer a Sume media or attachment URL for audio_url. Delete your own copies on the schedule your policy requires; this post does not cover consent or retention rules, and none of it is legal advice.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume