Post-call QA for a voice agent: transcribe 500 recorded calls for $40
Sume STT is batch, not live. Use it after the call: 500 eight-minute recordings cost about $40 a month at $0.01 a minute, with no speaker labels.

If you run a voice agent, Sume STT is not what listens during the call, but it can transcribe the recordings afterwards. At $0.01 per audio minute, 500 calls of eight minutes each is 4,000 minutes, or about $40 a month. Each job takes up to 10 minutes of audio, so an eight-minute call fits in one.
What it is good for
Batch transcription suits review work: spot checks for missed disclosures, finding calls where the caller asked for a human, building a searchable archive, and generating training examples. It does not suit anything that must happen mid-call, because jobs are asynchronous and the audio must already be a file at a public HTTPS URL.
| Calls per month | Length | Audio minutes | List cost at $0.01 a minute |
|---|---|---|---|
| 100 | 8 minutes | 800 | about $8 |
| 500 | 8 minutes | 4,000 | about $40 |
| 500 | 12 minutes (2 jobs each) | 6,000 | about $60 |
| 2,000 | 3 minutes | 6,000 | about $60 |
Two limits to plan around
First, Sume STT has no speaker labels. If your recording keeps the caller and the agent on separate tracks, transcribe each track as its own job and merge the transcripts by timestamp. A mixed mono file will come back as one stream of text. Second, set language_code when you know the language, and send duration_seconds so the admission estimate is close to the real cost.
- Calls over 10 minutes need splitting into chunks before submission.
- Send an idempotency key per call so a retry does not double-bill.
- Use sentence segmentation if you want time ranges for review tooling.
A sampling alternative
You may not need all 500. If the aim is quality assurance rather than a full archive, transcribe a random sample of 10 percent. That is 50 calls, 400 minutes and about $4 a month. Sample more heavily on calls where a signal fired, such as a long silence or a hang-up, and less on routine ones.
Keep the decision rule in code, so the sample is not chosen by whoever is reviewing.
Privacy and consent
Recordings often contain personal data. The audio has to be reachable by URL, so use a short-lived link where you can, delete the source when the review is done, and check your recording-consent and retention obligations in each market. That part is yours, not a setting on a Sume job.
Sources
Related posts
More in Use cases
- Turn a voice memo into a narrated short: transcribe, edit, re-voice
Speak your idea into a phone, transcribe it with Sume STT, tidy the text, then re-voice and caption it. The memo becomes the script, not the audio.
- YouTube watch series button: viewers may land on episode 7 first
A viewer who finds one Short in their feed sees a watch series button, so any episode can be the first they see. Write and render episodes that stand alone.
- Wedding save-the-date clip: first frame or reference photos
Send one engagement photo as a first frame to animate it, or up to 10 reference images to Omni Flash 1.1 to keep the couple's look. 8-second prices inside.
- What is a consistent story for a YouTube show? AI series checklist
YouTube asks shows for the same characters or hosts, one long story, or one topic. A checklist for keeping an AI-made series consistent across episodes.
Written by Sume