STT-ready audio: detach as 16000 Hz mono wav in one call

Sume audio detach can output the STT shape directly: sample_rate 16000 with channels mono, in sample-exact wav, for $0.01 per job.

5 min readSume
All posts

To get an audio file shaped for speech-to-text from a video, call Sume audio detach with sample_rate: 16000 and channels: "mono". The docs call that combination the STT shape. The default output format is sample-exact wav (pcm_s16le), the video is not changed, and the public rate is $0.01 per job. The new audio artifact gets its own artf_ id and a media.sume.com URL.

The request

The input must be a media.sume.com video in your workspace; import files first with POST /v1/media-imports. An Idempotency-Key is required, and the default mode is async.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: detach-stt-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "format": "wav",
    "channels": "mono",
    "sample_rate": 16000
  }'

Option table

Audio detach options, read 2026-10-08
FieldValuesDefault
formatwav (pcm_s16le, sample-exact) or mp3 (128 kbps)wav
range{ start, end? } in secondswhole track
channelssource or monosource
sample_rate16000, 44100 or 48000inherit the source

Limits and refusals

The source may be up to 1800 seconds and the output up to 900 seconds. A whole track longer than 900 seconds needs a range. A video with no audio track fails with detach_source_has_no_audio; probe has_audio first with video inspect.

Do not send ffmpeg fields such as af, filter, codec or cmd. The server compiles ffmpeg itself and rejects those with ffmpeg_fields_rejected.

Why shape the audio first

Sume's own transcript path, transcribe: true on video inspect, runs on a clip directly and does not need a detach step. Detach pays off when you want the audio file for another tool, such as an outside transcription service, or when you want to split one long track into ranges with timeline audio. For reference, Microsoft states its MAI-Transcribe-2-Streaming price as $0.54 per hour of audio through the end of the year (read 2026-10-08); a Sume detach job is $0.01 whatever the length, then your transcription service bills separately.

  • Keep wav if the file will be joined again or used for lip-sync.
  • Use mp3 only when file size matters.
  • Detach once, split many times.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume