Extract 16 kHz mono audio from a video for speech-to-text

Set sample_rate 16000 and channels mono on Sume's audio detach to get the speech-to-text shape from a video. Options, the 900 second cap and the price.

4 min readSume
All posts

To pull speech-to-text-ready audio out of a video, call POST /v1/audio-detach with sample_rate: 16000 and channels: "mono". The docs call that combination the STT shape. Leave format alone and you get sample-exact wav, which is also what timeline_create and the timeline audio endpoint take.

Everything here is from the Audio detach docs, read 2026-09-29.

Which options can I set?

Audio detach program fields, from the Audio detach docs, read 2026-09-29.
FieldValues
formatwav (default, pcm_s16le) or mp3 (128 kbps)
range{ start, end? } in seconds; whole track if omitted
channelssource (default) or mono
sample_rate16000, 44100 or 48000; inherits the source if omitted

What does the request look like?

The source must be a workspace media.sume.com video, and Idempotency-Key is required. The default mode is async, so poll GET /v1/jobs/:id/result for audio_url.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audio-detach-stt-001" \
  -d '{"video_url":"https://media.sume.com/artifacts/artf_demo/talk.mp4","channels":"mono","sample_rate":16000}'

How long a video can I detach?

Source up to 1800 seconds and output up to 900 seconds. A whole track past 900 seconds needs a range. For many ranges, detach once and split the track with timeline audio's split operation instead of detaching repeatedly.

Do I need to detach just to transcribe?

Not always. POST /v1/video-inspect with transcribe: true runs Sume's speech-to-text on the clip directly, at a listed $0.01 per audio minute. Detach is for when you want the audio file itself. It is listed at $0.01 per job; confirm in GET /v1/catalog.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume