Get speech audio from a video for transcription: 16 kHz mono, $0.11
Sume audio detach with channels mono and sample_rate 16000 gives the speech-to-text shape. A 10-minute video costs $0.01 to detach plus $0.10 to transcribe.

To pull transcription-ready audio out of a video, call POST /v1/audio-detach with channels: "mono" and sample_rate: 16000. The docs name that pairing as the speech-to-text shape. A 10-minute video costs $0.01 to detach and $0.10 to transcribe at $0.01 per audio minute, so $0.11 in total.
The request
Audio detach takes one media.sume.com video from your workspace and returns a new audio artifact; the video is untouched. The default output is sample-exact WAV (pcm_s16le). format: "mp3" gives 128 kbps instead.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-stt-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/interview.mp4",
"format": "wav",
"channels": "mono",
"sample_rate": 16000
}'The arithmetic
Detach is $0.01 per job whatever the length. Speech-to-text is $0.01 per audio minute (the catalog's estimate cap is 10 minutes, $0.10). For a 10-minute video: $0.01 + 10 x $0.01 = $0.11.
The output is capped at 900 seconds and the source at 1,800 seconds. A 20-minute video (1,200 seconds) therefore needs a range on each of two detach jobs, for example {start: 0, end: 600} and {start: 600}. Cost: 2 x $0.01 + 20 x $0.01 = $0.22. A whole track over 900 seconds with no range is refused.
| Video length | Detach jobs | STT minutes | Total |
|---|---|---|---|
| 5 min | 1 ($0.01) | 5 ($0.05) | $0.06 |
| 10 min | 1 ($0.01) | 10 ($0.10) | $0.11 |
| 20 min | 2 ($0.02) | 20 ($0.20) | $0.22 |
Check for audio first
A source without an audio track fails with detach_source_has_no_audio. To avoid the failed job, run video inspect with frames: false and read probe.has_audio. Also note that video inspect can transcribe on its own with transcribe: true, so detach is for when you want the audio file itself, for instance to feed lip-sync, a timeline spine, or a split.
Why 16 kHz mono
Speech-to-text models are usually run on mono audio at 16 kHz, and the Sume docs name that pairing as the STT shape. Sending it avoids carrying stereo and high sample rates through the next step. The default WAV keeps the source's channels and sample rate, so a 48 kHz stereo video gives a larger file than you need for transcription.
Other options are sample_rate: 44100 or 48000 for music or sync work, and range to take only part of the track, for example {start: 60, end: 120} for the second minute. If you omit end, the range is open-ended.
Choosing WAV or MP3
WAV is sample-exact and what the timeline render, the audio split, and speech-to-text expect, so it is the default. MP3 at 128 kbps is smaller but is not sample-exact. For a 10-minute transcription job the file size rarely matters, so stay with WAV and keep the option of joining or splitting the audio later without drift. A mono 16 kHz WAV is about 1.9 MB per minute by plain arithmetic (16,000 samples x 2 bytes x 60 seconds), so 10 minutes is about 19 MB.
Reading the result
The job is async by default. Poll GET /v1/jobs/:id/result for kind: audio_detach with audio_url, duration_seconds, format, channels, sample_rate, and source_duration_seconds. The sample rate is null when you omitted it, because the value then comes from the source.
Sources
Related posts
More in Media tools
- Dim an avatar clip to 0.45 for overlay text: run the free check first
Darken a Sume avatar video with video-filter dim 0.45, run the unbilled /check call first, then pay $0.02 for the encode. What dim does and what it refuses.
- Dim busy b-roll so captions read: video filter, then captions, $0.22
Darken a busy clip with video filter dim 0.45 for $0.02, then burn captions for $0.20. Ten clips cost $2.20. The dim check endpoint is free.
- Does the Kling motion control prompt change the movement? No
On Sume's Kling 3.0 Motion Control route the motion video drives all movement and the prompt only steers appearance. Fields, limits and $0.1575 per second.
- Eight inspect stills on a 45 s Snap ad: one lands in the first 6 s
Video inspect's default eight stills on 45 seconds sit 5.6 s apart; only 2.8 s falls in Snap's non-skippable first 6. Use frames at 1, 3 and 5.9 instead.
Written by Sume