Transcribe 1,000 twenty-second clips: pack 20 per file, then run STT
1,000 clips of 20 s cost $10.00 on Sume STT if billed one minute each, but $4.00 if you pack 20 clips per file with Timeline audio.

Run 1,000 clips of 20 seconds through Sume STT one by one and the bill is up to $10.00, because the minimum is one second but short clips still round to the audio minute. Pack 20 clips per file with the timeline audio concat endpoint and it drops to about $4.00; the batch is 400 seconds, under the 10-minute STT cap.
Arithmetic
Sume STT 1.0 is $0.01 per audio minute with a minimum of 1 second and a 10-minute cap per request. I treat each 20-second clip as costing a full minute, so the worst case is 1,000 x $0.01 = $10.00. Packed 20 per file, each batch is 400 seconds, 7 minutes when rounded up: $0.07 for STT plus $0.01 for the concat job, $0.08 per batch, times 50 batches is $4.00. Microsoft's announcements give MAI-Transcribe-2 as $0.10 per hour limited-time through end of 2026; 1,000 clips are 20,000 seconds or 5.56 hours, so $0.56.
| Path | Basis | Cost |
|---|---|---|
| Sume STT, one job per clip | 1,000 x $0.01, one minute each | $10.00 |
| Sume concat 20 clips then STT | 50 x ($0.01 + $0.07) | $4.00 |
| MAI-Transcribe-2 | 5.56 h x $0.10, limited-time rate | $0.56 |
| MAI-Transcribe-2-Streaming | 5.56 h x $0.54, intro rate | $3.00 |
Packing with the audio endpoint
POST /v1/timeline-1.0/audio joins parts in order for a flat $0.01 and returns segments[] offsets, so you can map the transcript words back to each source clip. It accepts 1 to 20 parts and returns media.sume.com URLs. Check the docs page for the exact request body before you write the loop; the example only shows the final STT call on the packed file.
curl -sS -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: pack-001" \
-d '{"audio_url":"https://media.sume.com/packs/pack-001.mp3","language_code":"en","duration_seconds":400}' \
| jq -r '.data.request_id'
Mapping words back
Word timings are always returned. Subtract each clip's offset from the concat response's segments[] to place words in the original clip. Add a second of silence between clips if boundary words matter.
Where this stops
Microsoft's MAI-Transcribe-2 prices are limited-time rates, and Sume does not document a keyword biasing field. I did not test accuracy for either. Packing only helps when clips do not need independent language hints.
Sources
Related posts
More in Developers
- Transcribe a 9-minute recording in two calls for ten cents
Detach the audio at 16 kHz mono ($0.01), then send it to Sume STT ($0.01 a minute): nine minutes of speech to text for $0.10 in two requests.
- Transcribe a MAI-Voice-2.1 clip with Sume STT: Python audio_url run
Host a MAI-Voice-2.1 or Flash clip at a public HTTPS URL and Sume STT returns text, word times and sentences for 1 cent a minute. A 25-line Python run.
- Translate a pack into 8 languages in parallel: queue limits by plan
Eight Ideogram 4.5 edits fit the Pro queue (24 accepted jobs) but not Free (6), so 2 get 429 queue_full. Python thread pool with retry; cost is $0.60 at medium.
- Transparent AI image: PNG or WebP, not JPEG, on GPT Image 2.5
For a transparent AI image, request png or webp with background transparent on GPT Image 2.5. JPEG has no alpha. Code to request and verify the alpha channel.
Written by Sume