Audio detach limits: 1800 s source, 900 s output, wav or mp3
Sume audio detach turns one hosted video into a wav or mp3 for $0.01. Source up to 1800 s, output up to 900 s, with the range, channel and sample rate options.

Audio detach extracts the audio track of one Sume-hosted video as a new wav (default) or mp3 for $0.01 a job. The source video may run up to 1800 seconds and the output up to 900 seconds, so a track longer than 900 seconds needs a range. The video itself is untouched. Details are in Audio detach.
The options
| Field | Values | Default |
|---|---|---|
format | wav (pcm_s16le, sample-exact) or mp3 (128 kbps) | wav |
range | { start, end? } in seconds | whole track |
channels | source or mono | source |
sample_rate | 16000, 44100 or 48000 | inherits the source |
Which format
Wav is the format other Sume steps read: the timeline audio input, the Timeline audio route and speech to text all use it. Choose mp3 only when you are handing the file to a person or a player and file size matters. For transcription, pair 16000 with mono; the docs call that the STT shape.
A request and its guardrails
The source must already be on media.sume.com in your workspace. The server will not fetch from the open internet, so import first with POST /v1/media-imports. An Idempotency-Key is required.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-mp3-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "mp3",
"range": { "start": 60, "end": 360 }
}'Cost of splitting a long file
A 25-minute recording needs two detach jobs of at most 900 seconds each, so $0.02. If you want many ranges from one track, detach once and split with Timeline audio instead of paying per range.
Sending ffmpeg-style fields such as filter or codec is rejected with ffmpeg_fields_rejected; the server builds the command itself.
Refusals to expect
Before detaching, probe the clip with video inspect and check probe.has_audio (frames: false is enough); a silent source fails rather than returning an empty file. Plan the range from the probed duration so the request never trips the codes below.
| Code | When |
|---|---|
| audio_detach_range_empty | range.end is at or before range.start, or the range is longer than 900 s |
| detach_source_has_no_audio | The source has no audio track |
| detach_start_past_source | range.start is past the probed duration |
| source_duration_exceeded | The source is longer than 1800 s |
Sources
Related posts
More in Media tools
- Captions from a video's speech: $0.20 a job, no SRT needed
Sume's video-captions endpoint transcribes the speech and burns styled captions for $0.20 a job on clips up to 60 s. Optional script_text corrects the wording.
- Check an AI voiceover by transcribing it: a script diff for 4 cents
Run Sume STT on a finished TTS take and diff it against the script to catch misread numbers and names. 4 cents per 450-character line, TTS plus the check.
- Trim and conform to 1080x1920 at 30 fps in one video-trim job
Video trim's output field re-encodes to a set width, height and 24/25/30/60 fps in the same $0.02 job, exact precision only. Request, ranges and refusal codes.
- Crop a 21:9 or 2.39:1 film clip to 9:16 with video filter
Width fractions for cropping ultrawide 21:9, 2.39:1 and 32:9 clips to 9:16 with one Sume video-filter crop op, plus the 0.05 minimum side and a free check.
Written by Sume