A 28-minute talk transcribed on Sume: 3 detach ranges plus STT = $0.31
Audio detach caps output at 900 s and STT estimates cap at 10 minutes. Here is how a 28-minute talk splits into three ranges, and what it costs: $0.31.

A 28-minute talk (1,680 seconds) fits Sume's 1,800-second audio-detach source cap, but not its 900-second output cap, and Sume STT reserves for at most 10 minutes per job. Cut it into three ranges of 600, 600 and 480 seconds, detach each as 16 kHz mono wav, and send each file to STT. The bill is 3 x $0.01 for detach plus 28 minutes x $0.01 for STT = $0.31.
Why three ranges
The audio-detach docs set the caps: source up to 1,800 seconds, output up to 900 seconds, so a whole track longer than 900 seconds needs a range. The STT catalog entry says Sume reserves from audio-minute estimates and that omitting duration_seconds reserves one minute, with a maximum of 10 minutes. The request schema accepts duration_seconds from 1 to 600. Ranges of 600 seconds keep every job inside that estimate.
| Step | Input | Cost |
|---|---|---|
| Detach range 0-600 | 1 job, wav, 16000 Hz, mono | $0.01 |
| Detach range 600-1200 | 1 job | $0.01 |
| Detach range 1200-1680 | 1 job | $0.01 |
| STT on 600 s + 600 s + 480 s | 28 audio minutes x $0.01 | $0.28 |
| Total | $0.31 |
The detach request
Import the video first with POST /v1/media-imports, because detach reads only this workspace's media.sume.com files and rejects off-host URLs at admit. Then send one request per range, each with its own Idempotency-Key so a retried range is not billed twice. 16000 with channels: "mono" is the speech-to-text shape named in the docs.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: talk-part-2" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"range": { "start": 600, "end": 1200 },
"channels": "mono",
"sample_rate": 16000
}'What fails, and what each error means
Detach has stable refusal codes, so a script can branch on them instead of parsing text. source_duration_exceeded means the source is longer than 1,800 seconds, and a 31-minute talk would hit it; trim it first or detach only ranges you need. audio_detach_range_empty means the range end is not after its start, or the range is longer than 900 seconds. detach_start_past_source means a range starts after the probed duration, which is what a hard-coded fourth range would do on this 28-minute file.
A source with no audio track fails with detach_source_has_no_audio. The docs suggest checking probe.has_audio with video inspect first; frames: false is enough, so the check does not need stills.
Stitching the transcripts
Each STT result carries words[] with { word, start, end } in seconds from the start of that audio file. After the second range, add 600 to every timestamp; after the third, add 1,200. Keep the offsets in your job metadata so a retry does not lose them. The metadata field on the STT request is stored on the job for exactly this kind of bookkeeping, for example the part number and its offset in seconds.
Microsoft's MAI-Transcribe-2 page lists 60 languages, a streaming variant, diarisation, timestamps and keyword biasing. Sume STT takes a language_code hint and has no vocabulary field, so a speaker's name that the model mis-hears has to be fixed in the text afterwards. The page does not list a price, so this post does not compare prices.
Sources
Related posts
More in Developers
- 34 pause cuts, 20 audio parts: split the spine in two levels
Timeline audio.parts holds 20 slices. For 34 pause cuts, build two concat files of 17, feed them as two parts, and re-base each video slot start.
- 37 video jobs submitted at once: which Sume plan accepts them all
Free accepts 6 of 37 video jobs, Pro 24, Startup all 37 with 11 spare slots, Scale all 37. Concurrency, queue and accepted-capacity table from the Sume docs.
- 4 of Sume's 17 aspect ratios exceed 3:1, GPT Image 2.5's size cap
Sume normalizes 17 aspect ratios; 1:4, 4:1, 1:8 and 8:1 are wider than the 3:1 limit on GPT Image 2.5 custom sizes. A check script and the 2 other rules.
- 42 voice lines and the 20-part concat cap: four joins for 4 cents
Sume timeline audio joins at most 20 parts per job at $0.01 each. For 42 voice lines run three joins of 20, 20 and 2, then a final join, four jobs and $0.04.
Written by Sume