QA 1,000 voice-agent calls with speech-to-text on Sume
Transcribe call recordings to review a voice agent. Sume STT is $0.01 per audio minute with a 10-minute reservation limit: the math for 1,000 calls.

Transcribing 1,000 six-minute call recordings with Sume STT costs about $60 at the public rate of $0.01 per audio minute, and each job reserves for up to 10 minutes. It gives you text and word timings to review a voice agent such as Decagon Voice 3 after the fact.
What you get per call
stt_create takes a public HTTPS audio_url and returns text, language code, words with start and end times, and optional sentence segments. Pass duration_seconds (1 to 600) so the reservation matches the call, and a language_code hint such as en if you know it. Without a duration, Sume reserves one minute.
Sume's STT job returns word timings for every call, so you can flag silences, long gaps before an answer, or the exact second a caller interrupted.
The budget
Using the rate card figure of $0.01 per audio minute:
| Calls | Average length | Audio minutes | Cost at $0.01 per minute |
|---|---|---|---|
| 100 | 6 min | 600 | $6.00 |
| 1,000 | 6 min | 6,000 | $60.00 |
| 1,000 | 3 min | 3,000 | $30.00 |
| 5,000 | 6 min | 30,000 | $300.00 |
Limits to plan around
A single job reserves a maximum of 10 minutes, so a longer call must be split first. Audio detach extracts audio from a video as WAV or MP3 and can output 16 kHz mono, which is the speech-to-text shape, for $0.01 per job. Its output is capped at 900 seconds, and a longer track needs a range.
Calls must be reachable on a public HTTPS URL, preferably on the Sume media host. Import private files first.
Sample before you scale
You rarely need to read all 1,000. Transcribe a random 5 percent, score it against a checklist (greeting, verification, resolution, handoff), and only expand if the sample shows problems. Use dry_run to get the cost before each batch, and max_spend_usd to cap it.
Sume's STT 1.0 fixes diarization and audio-event tagging server-side and rejects diarize or tag_audio_events in the request, so you cannot ask for speaker labels. If you need to tell the agent from the caller, keep them on separate tracks in your recording.
Sources
Related posts
More in Developers
- queue_full 429 on a Sume submit: the reservation is released
A 429 queue_full releases or refunds the failed admission's reservation. Check refunded_usd_micros in /v1/usage, then retry with the same Idempotency-Key.
- Can I submit 100 AI video jobs at once? Queue limits by plan
Accepted capacity is slots plus queue: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Submit 100 at once and 94, 76, 52 or 0 get 429 queue_full.
- R httr2: POST /v1/images on Sume without throwing on a 502
Call Sume's image API from R with httr2: bearer key from the environment, a 40-second timeout, and req_error so a 502 body is readable.
- Re-render a prompts.txt of old Sora prompts with curl, jq and xargs
Turn a file of saved Sora prompts into Sume video jobs with four parallel curl calls, an idempotency key per prompt, then poll jobs.tsv and download each mp4.
Written by Sume