Transcribing a 2-hour call on Sume: 12 STT chunks, about $1.20

Sume STT takes 10 minutes per job at $0.01 a minute, so a 2-hour recording is 12 chunks and about $1.20 at list. How to cut, run, and re-stitch it.

4 min readSume
All posts

A two-hour recording is 7,200 seconds, and Sume STT accepts at most 600 seconds per job. That is 12 jobs, and at $0.01 per audio minute the list cost is about $1.20 for 120 minutes. The work is in cutting cleanly and putting the transcripts back in order.

The plan

Cut the audio into 10-minute pieces with a cut or detach step, submit each piece as its own speech-to-text job, then join the results by chunk order. Overlap helps you avoid losing a word at a seam, but each overlap minute is billed twice, so keep it short.

Sume rates and limits, read 2026-10-06 from the repo catalog and schemas
RecordingChunks of up to 10 minutesList cost at $0.01 a minute
45 minutes5about 45 cents
90 minutes9about 90 cents
120 minutes12about $1.20
120 minutes with 10-second overlaps12 chunks, 11 overlapsabout $1.20 plus 11 x 10 seconds, roughly 2 cents more

Stitching the transcripts

Keep an offset for each chunk. Chunk 3 starts at 1,200 seconds, so add 1,200 to every timestamp in its result before you merge. If you used overlaps, drop the first sentence of each chunk after the first when it repeats the end of the previous one. Sentence segmentation, which Sume STT supports with mode sentence, gives you time ranges to compare.

offsets = [i * 600 for i in range(12)]
for i, off in enumerate(offsets):
    print(f'chunk {i+1}: starts at {off}s')

A checklist before you start

Confirm each piece is at most 600 seconds, with a margin: cut at 595 seconds if your tool rounds up. Give each chunk a stable name with its index, so a retry of chunk 7 cannot be confused with chunk 8. Send an idempotency key derived from the file and the chunk index, so a network retry does not create a second job.

If a chunk fails, resubmit only that chunk. The other eleven transcripts stay valid, and you pay for the failed one only if it completed.

Where this does not fit

Sume STT has no speaker labels and no streaming. For a call where you need who said what, record each side on its own track and transcribe them separately, then merge by time. For a live call, use a streaming service during the call and run Sume STT afterwards on the recording if you want a second pass.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume