QA 1,000 voice-agent calls with speech-to-text on Sume

Transcribe call recordings to review a voice agent. Sume STT is $0.01 per audio minute with a 10-minute reservation limit: the math for 1,000 calls.

5 min readSume
All posts

Transcribing 1,000 six-minute call recordings with Sume STT costs about $60 at the public rate of $0.01 per audio minute, and each job reserves for up to 10 minutes. It gives you text and word timings to review a voice agent such as Decagon Voice 3 after the fact.

What you get per call

stt_create takes a public HTTPS audio_url and returns text, language code, words with start and end times, and optional sentence segments. Pass duration_seconds (1 to 600) so the reservation matches the call, and a language_code hint such as en if you know it. Without a duration, Sume reserves one minute.

Sume's STT job returns word timings for every call, so you can flag silences, long gaps before an answer, or the exact second a caller interrupted.

The budget

Using the rate card figure of $0.01 per audio minute:

Cost of a QA sample (read 2026-10-07)
CallsAverage lengthAudio minutesCost at $0.01 per minute
1006 min600$6.00
1,0006 min6,000$60.00
1,0003 min3,000$30.00
5,0006 min30,000$300.00

Limits to plan around

A single job reserves a maximum of 10 minutes, so a longer call must be split first. Audio detach extracts audio from a video as WAV or MP3 and can output 16 kHz mono, which is the speech-to-text shape, for $0.01 per job. Its output is capped at 900 seconds, and a longer track needs a range.

Calls must be reachable on a public HTTPS URL, preferably on the Sume media host. Import private files first.

Sample before you scale

You rarely need to read all 1,000. Transcribe a random 5 percent, score it against a checklist (greeting, verification, resolution, handoff), and only expand if the sample shows problems. Use dry_run to get the cost before each batch, and max_spend_usd to cap it.

Sume's STT 1.0 fixes diarization and audio-event tagging server-side and rejects diarize or tag_audio_events in the request, so you cannot ask for speaker labels. If you need to tell the agent from the caller, keep them on separate tracks in your recording.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume