STT-ready audio: detach as 16000 Hz mono wav in one call
Sume audio detach can output the STT shape directly: sample_rate 16000 with channels mono, in sample-exact wav, for $0.01 per job.

To get an audio file shaped for speech-to-text from a video, call Sume audio detach with sample_rate: 16000 and channels: "mono". The docs call that combination the STT shape. The default output format is sample-exact wav (pcm_s16le), the video is not changed, and the public rate is $0.01 per job. The new audio artifact gets its own artf_ id and a media.sume.com URL.
The request
The input must be a media.sume.com video in your workspace; import files first with POST /v1/media-imports. An Idempotency-Key is required, and the default mode is async.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-stt-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "wav",
"channels": "mono",
"sample_rate": 16000
}'Option table
| Field | Values | Default |
|---|---|---|
format | wav (pcm_s16le, sample-exact) or mp3 (128 kbps) | wav |
range | { start, end? } in seconds | whole track |
channels | source or mono | source |
sample_rate | 16000, 44100 or 48000 | inherit the source |
Limits and refusals
The source may be up to 1800 seconds and the output up to 900 seconds. A whole track longer than 900 seconds needs a range. A video with no audio track fails with detach_source_has_no_audio; probe has_audio first with video inspect.
Do not send ffmpeg fields such as af, filter, codec or cmd. The server compiles ffmpeg itself and rejects those with ffmpeg_fields_rejected.
Why shape the audio first
Sume's own transcript path, transcribe: true on video inspect, runs on a clip directly and does not need a detach step. Detach pays off when you want the audio file for another tool, such as an outside transcription service, or when you want to split one long track into ranges with timeline audio. For reference, Microsoft states its MAI-Transcribe-2-Streaming price as $0.54 per hour of audio through the end of the year (read 2026-10-08); a Sume detach job is $0.01 whatever the length, then your transcription service bills separately.
- Keep wav if the file will be joined again or used for lip-sync.
- Use mp3 only when file size matters.
- Detach once, split many times.
Sources
Related posts
More in Media tools
- Three 61-second renders cost $0.60; one 183-second render costs $0.40
Timeline 1.0 rounds each render up to a whole minute at $0.10. Three 61 s renders bill 6 minutes; one 183 s render bills 4. A YouTube Shorts length check too.
- TikTok 9:16 ad at 540x960: Timeline render plus captions is $0.30
TikTok lists 9:16 with a 540x960 minimum for non-Spark ads. A Sume Timeline render of up to 60 s ($0.10) plus a caption job ($0.20) is $0.30. Checks to run.
- A 17.3 s dance source on Genjutsu bills 18 s: $7.16 at 480p
A 17.3 s TikTok dance bills as 18 s on Genjutsu Motion Transfer: $7.16 at 480p, $15.33 at 720p. The arithmetic and a check before you pay for 720p.
- TikTok safe zone captions: anchor_ratio and safe_width_ratio on Sume
TikTok's ad page says safe zones vary by orientation and caption length and ships reference files. Sume's caption design fields move the line; ranges listed.
Written by Sume