Transcribing a 2-hour call on Sume: 12 STT chunks, about $1.20
Sume STT takes 10 minutes per job at $0.01 a minute, so a 2-hour recording is 12 chunks and about $1.20 at list. How to cut, run, and re-stitch it.

A two-hour recording is 7,200 seconds, and Sume STT accepts at most 600 seconds per job. That is 12 jobs, and at $0.01 per audio minute the list cost is about $1.20 for 120 minutes. The work is in cutting cleanly and putting the transcripts back in order.
The plan
Cut the audio into 10-minute pieces with a cut or detach step, submit each piece as its own speech-to-text job, then join the results by chunk order. Overlap helps you avoid losing a word at a seam, but each overlap minute is billed twice, so keep it short.
| Recording | Chunks of up to 10 minutes | List cost at $0.01 a minute |
|---|---|---|
| 45 minutes | 5 | about 45 cents |
| 90 minutes | 9 | about 90 cents |
| 120 minutes | 12 | about $1.20 |
| 120 minutes with 10-second overlaps | 12 chunks, 11 overlaps | about $1.20 plus 11 x 10 seconds, roughly 2 cents more |
Stitching the transcripts
Keep an offset for each chunk. Chunk 3 starts at 1,200 seconds, so add 1,200 to every timestamp in its result before you merge. If you used overlaps, drop the first sentence of each chunk after the first when it repeats the end of the previous one. Sentence segmentation, which Sume STT supports with mode sentence, gives you time ranges to compare.
offsets = [i * 600 for i in range(12)]
for i, off in enumerate(offsets):
print(f'chunk {i+1}: starts at {off}s')A checklist before you start
Confirm each piece is at most 600 seconds, with a margin: cut at 595 seconds if your tool rounds up. Give each chunk a stable name with its index, so a retry of chunk 7 cannot be confused with chunk 8. Send an idempotency key derived from the file and the chunk index, so a network retry does not create a second job.
If a chunk fails, resubmit only that chunk. The other eleven transcripts stay valid, and you pay for the failed one only if it completed.
Where this does not fit
Sume STT has no speaker labels and no streaming. For a call where you need who said what, record each side on its own track and transcribe them separately, then merge by time. For a live call, use a streaming service during the call and run Sume STT afterwards on the recording if you want a second pass.
Sources
Related posts
More in Use cases
- 2-minute AI explainer at 1080p: Veo 3.1 Lite $9.60, Standard $48
120 seconds of 1080p video costs $9.60 on Veo 3.1 Lite, $14.40 on Fast, $48 on Standard and $22.50 on Omni Flash on Sume. 4K totals included. Read 2026-10-06.
- Two voices, one conversation: Sume TTS jobs joined by concat
Make a two-person dialogue file with Sume: one TTS job per line using two avatar voices, then one Timeline audio concat. Cost, code and the gap.
- 18-second UGC-style ad with a product image: $3.49, $4.64 or $10.44
An 18-second talking-avatar UGC ad with product_image costs $3.49 on standard, $4.64 on plus and $10.44 on max. The no-product rates and when to pay more.
- UGC-style ad: test five hooks on one body with Timeline plans
Join five 3-second hooks to one 12-second body clip with Timeline 1.0. Plan each cut unbilled, then render the winners at $0.10 a minute.
Written by Sume