Post-call QA for a voice agent: transcribe 500 recorded calls for $40

Sume STT is batch, not live. Use it after the call: 500 eight-minute recordings cost about $40 a month at $0.01 a minute, with no speaker labels.

4 min readSume
All posts

If you run a voice agent, Sume STT is not what listens during the call, but it can transcribe the recordings afterwards. At $0.01 per audio minute, 500 calls of eight minutes each is 4,000 minutes, or about $40 a month. Each job takes up to 10 minutes of audio, so an eight-minute call fits in one.

What it is good for

Batch transcription suits review work: spot checks for missed disclosures, finding calls where the caller asked for a human, building a searchable archive, and generating training examples. It does not suit anything that must happen mid-call, because jobs are asynchronous and the audio must already be a file at a public HTTPS URL.

Sume rates and limits, read 2026-10-06 from the repo catalog and schemas
Calls per monthLengthAudio minutesList cost at $0.01 a minute
1008 minutes800about $8
5008 minutes4,000about $40
50012 minutes (2 jobs each)6,000about $60
2,0003 minutes6,000about $60

Two limits to plan around

First, Sume STT has no speaker labels. If your recording keeps the caller and the agent on separate tracks, transcribe each track as its own job and merge the transcripts by timestamp. A mixed mono file will come back as one stream of text. Second, set language_code when you know the language, and send duration_seconds so the admission estimate is close to the real cost.

  • Calls over 10 minutes need splitting into chunks before submission.
  • Send an idempotency key per call so a retry does not double-bill.
  • Use sentence segmentation if you want time ranges for review tooling.

A sampling alternative

You may not need all 500. If the aim is quality assurance rather than a full archive, transcribe a random sample of 10 percent. That is 50 calls, 400 minutes and about $4 a month. Sample more heavily on calls where a signal fired, such as a long silence or a hang-up, and less on routine ones.

Keep the decision rule in code, so the sample is not chosen by whoever is reviewing.

Privacy and consent

Recordings often contain personal data. The audio has to be reachable by URL, so use a short-lived link where you can, delete the source when the review is done, and check your recording-consent and retention obligations in each market. That part is yours, not a setting on a Sume job.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume