Transcribe a 9-minute recording in two calls for ten cents
Detach the audio at 16 kHz mono ($0.01), then send it to Sume STT ($0.01 a minute): nine minutes of speech to text for $0.10 in two requests.

A 9-minute video on Sume media storage costs $0.10 to turn into a transcript with word timings: one audio detach job at $0.01, then one STT job at $0.01 per audio minute, which is $0.09 for 9 minutes. Detach first with format: wav, channels: mono and sample_rate: 16000, which the audio detach page calls the STT shape.
Why two calls
Sume STT takes an audio_url, not a video. The audio detach page says its default sample-exact wav is the format that POST /v1/stt-1.0/transcribe uses. Detach reads one media.sume.com video that your workspace owns, so import any outside file first with POST /v1/media-imports. The video is untouched and you get a new audio artifact.
Microsoft's MAI-Transcribe-2 page promises timestamps and diarisation among its features (read 2026-10-05). On Sume STT, word timings are always returned and speaker options are fixed server-side, so plan for words, not speaker labels.
Call one: detach
Detach is async by default. With mode: "sync" the API waits up to 30 seconds for a completed job, and otherwise returns 202 so you poll GET /v1/jobs/:id/status. An Idempotency-Key is required.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-rec-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "wav",
"channels": "mono",
"sample_rate": 16000,
"mode": "sync"
}'Call two: transcribe
Take audio_url from the detach result, and add duration_seconds so the reservation matches the real length. Without it Sume reserves one minute, and the maximum hint is 10 minutes. Add segmentation: { mode: "sentence" } when you want caption-sized lines as well as words.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-rec-001" \
-d '{
"audio_url": "AUDIO_URL_FROM_DETACH_RESULT",
"duration_seconds": 540,
"segmentation": { "mode": "sentence" }
}'The receipt
Detach is $0.01 per job, run on worker ffmpeg with no provider inference. STT is billed per ceil audio minute of the hint.
| Recording | Detach | STT at $0.01 a minute | Total |
|---|---|---|---|
| 3 minutes | $0.01 | $0.03 | $0.04 |
| 9 minutes | $0.01 | $0.09 | $0.10 |
| 10 minutes | $0.01 | $0.10 | $0.11 |
Limits that shape the recipe
If you only need one call, video inspect can transcribe in place with transcribe: true. It bills its own Modal compute on top of the same $0.01 a minute, so two calls are easier to price exactly.
- Detach source is capped at 1,800 seconds and its output at 900 seconds, so anything over 15 minutes needs a
range. - The STT duration hint tops out at 10 minutes, so chunk longer audio and transcribe each piece.
- A video without an audio track fails with
detach_source_has_no_audio. Probe first with video inspect andframes: false.
What to store
Keep the wav from the detach job and the word list from the STT job together. The wav is the input to any later transcription, and the words feed captions: Sume's caption job accepts words directly and skips speech-to-text, so a transcript you already paid for does not have to be paid for again when you decide to burn subtitles.
If a fix is needed, edit the text, not the audio. Correct the words, resend them to captions, and leave the detach artifact alone. The only reason to run detach twice is a different range or a different sample rate.
Common mistakes
- Sending a video URL to STT. It wants audio, so detach first.
- Forgetting
duration_seconds, which leaves the reservation at one minute for a nine-minute file. - Reusing an
Idempotency-Keywith a different body, which is rejected instead of creating a second job. - Importing a file from outside
media.sume.comand skippingPOST /v1/media-imports.
When the recording is longer
For a 25-minute recording, detach in ranges that each stay under the 900-second output cap, then transcribe each piece with its own duration hint. Three detach jobs and three STT jobs cost 3 cents plus 25 cents, so 28 cents, and you stitch the word lists together by adding each range's start offset to its timings. Keep the offsets in your own table; the transcript itself does not know where its piece began.
Sources
Related posts
More in Developers
- Transcribe a MAI-Voice-2.1 clip with Sume STT: Python audio_url run
Host a MAI-Voice-2.1 or Flash clip at a public HTTPS URL and Sume STT returns text, word times and sentences for 1 cent a minute. A 25-line Python run.
- Translate a pack into 8 languages in parallel: queue limits by plan
Eight Ideogram 4.5 edits fit the Pro queue (24 accepted jobs) but not Free (6), so 2 get 429 queue_full. Python thread pool with retry; cost is $0.60 at medium.
- Transparent AI image: PNG or WebP, not JPEG, on GPT Image 2.5
For a transparent AI image, request png or webp with background transparent on GPT Image 2.5. JPEG has no alpha. Code to request and verify the alpha channel.
- Trim an AI-generated clip without a media import
Generated artifacts already live on media.sume.com, which is the host video-trim accepts. Use media-imports only for clips from outside Sume.
Written by Sume