Review voice agent call recordings: STT word timings on Sume
After you ship a Gemini Live or other voice agent, transcribe the recordings with Sume STT: word timings, sentence segments, a 10-minute cap per request.

Send each recording to Sume's STT 1.0 route as a public HTTPS audio_url and read back text plus word-level timings. Add segmentation mode sentence and you also get sentence segments, which map cleanly onto turns of a call. A request covers up to 10 minutes of audio, because duration_seconds tops out at 600, so split longer calls first.
Google's Gemini API changelog (read 2026-10-02) lists Gemini 3.8 Live GA on 2026-09-15, so more teams are shipping voice agents now and then needing to read what was said. STT is a batch job, not part of the live session.
What does the request look like?
audio_url is required. language_code is optional, for example en or ko; omit it for auto-detect. duration_seconds from 1 to 600 improves the usage reservation; omit it and Sume reserves 1 minute. Word timings are always returned, so there is no flag to ask for them.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: call-0142" \
-d '{"audio_url": "https://media.sume.com/artifacts/example/call.wav", "duration_seconds": 240, "segmentation": {"mode": "sentence"}}'What do I get back, and what do I not?
The completed result carries text, language fields when available, and words with start and end seconds from the audio start. The table lists what the docs I read show and what they do not.
| Item | Present | Note |
|---|---|---|
| Full text | Yes | Public-safe text |
| Word timings | Yes | word, start, end in seconds |
| Sentence segments | On request | segmentation.mode sentence |
| Speaker labels | No field in the docs I read | Provider knobs such as diarization are fixed server-side |
How do I handle a 25-minute call?
Split it into pieces of 10 minutes or less with Timeline audio split, which takes a top-level url and 1 to 20 ranges and returns durable media.sume.com files. Transcribe each piece, then add each piece's start offset to its word times.
What does it cost?
The video-inspect page lists the public STT rate as $0.01 per audio minute and says to confirm it live in GET /v1/catalog.
What should I do with the transcript?
Start with the failures. Search each transcript for the phrases your agent should never say and for the points where the caller repeats themselves; the word timings let you jump to that second of the recording. Keep the job id and the original recording URL with the transcript so a reviewer can listen to the source.
Treat the text as a working record, not a verdict. Check a sample against the audio before you build metrics on it, particularly for names, numbers and mixed-language calls, and set language_code when you know the language.
How do I keep reviewers out of private audio?
Only send recordings you are allowed to process, and use a URL that is reachable by Sume when the job runs. Sume's docs prefer a Sume media or attachment URL for audio_url. Delete your own copies on the schedule your policy requires; this post does not cover consent or retention rules, and none of it is legal advice.
Sources
Related posts
More in Developers
- Voice API deadlines, October 2026 to February 2027
A calendar of voice and transcription API changes from vendor pages: Gemini TTS price rise, OpenAI transcription shutdown, and the xAI voice alias move.
- Wait for an avatar video job with the SDK: waitForJob and its timeout
Avatar jobs need waitForJob, not waitForRun. Defaults, the 20-minute timeout, and why a SumeJobTimeoutError neither cancels nor refunds the render.
- waitForRun and 429 or 503 on a status read: it keeps polling
A 429 or 5xx on a status read does not fail the run. waitForRun tolerates 6 consecutive transient read failures and keeps waiting. Options and what to log.
- Wan 2.2 TI2V-5B in Diffusers: 121 frames, 24 fps, 24GB GPU
A working Diffusers example for Wan2.2-TI2V-5B: 704x1280 portrait frames, 121 frames at 24 fps, 24GB VRAM minimum, and when to use a hosted route instead.
Written by Sume