Transcribe 1,000 twenty-second clips: pack 20 per file, then run STT

1,000 clips of 20 s cost $10.00 on Sume STT if billed one minute each, but $4.00 if you pack 20 clips per file with Timeline audio.

4 min readSume
All posts

Run 1,000 clips of 20 seconds through Sume STT one by one and the bill is up to $10.00, because the minimum is one second but short clips still round to the audio minute. Pack 20 clips per file with the timeline audio concat endpoint and it drops to about $4.00; the batch is 400 seconds, under the 10-minute STT cap.

Arithmetic

Sume STT 1.0 is $0.01 per audio minute with a minimum of 1 second and a 10-minute cap per request. I treat each 20-second clip as costing a full minute, so the worst case is 1,000 x $0.01 = $10.00. Packed 20 per file, each batch is 400 seconds, 7 minutes when rounded up: $0.07 for STT plus $0.01 for the concat job, $0.08 per batch, times 50 batches is $4.00. Microsoft's announcements give MAI-Transcribe-2 as $0.10 per hour limited-time through end of 2026; 1,000 clips are 20,000 seconds or 5.56 hours, so $0.56.

1,000 clips of 20 seconds, 5.56 hours total (read 2026-10-05)
PathBasisCost
Sume STT, one job per clip1,000 x $0.01, one minute each$10.00
Sume concat 20 clips then STT50 x ($0.01 + $0.07)$4.00
MAI-Transcribe-25.56 h x $0.10, limited-time rate$0.56
MAI-Transcribe-2-Streaming5.56 h x $0.54, intro rate$3.00

Packing with the audio endpoint

POST /v1/timeline-1.0/audio joins parts in order for a flat $0.01 and returns segments[] offsets, so you can map the transcript words back to each source clip. It accepts 1 to 20 parts and returns media.sume.com URLs. Check the docs page for the exact request body before you write the loop; the example only shows the final STT call on the packed file.

curl -sS -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: pack-001" \
  -d '{"audio_url":"https://media.sume.com/packs/pack-001.mp3","language_code":"en","duration_seconds":400}' \
  | jq -r '.data.request_id'

Mapping words back

Word timings are always returned. Subtract each clip's offset from the concat response's segments[] to place words in the original clip. Add a second of silence between clips if boundary words matter.

Where this stops

Microsoft's MAI-Transcribe-2 prices are limited-time rates, and Sume does not document a keyword biasing field. I did not test accuracy for either. Packing only helps when clips do not need independent language hints.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume