ffmpeg -vn -ac 1 -ar 16000 as audio detach API fields
Audio detach rejects raw ffmpeg fields such as af, filter, cmd and codec. Map -ac, -ar and -ss/-to to channels, sample_rate and range; mp3 is 128 kbps.

Audio detach compiles ffmpeg on the server, so sending af, filter, ffmpeg, cmd or codec fails with ffmpeg_fields_rejected. The four things people set on the command line have fields: channels, sample_rate, range and format.
Which flag maps to which field?
The mapping below follows the Audio detach docs; the flag column is the usual ffmpeg meaning.
| ffmpeg idea | Detach field |
|---|---|
Drop the video (-vn) | Implicit: the result is a new audio artifact |
Mono (-ac 1) | channels: "mono" (default source) |
Sample rate (-ar 16000) | sample_rate: 16000, 44100 or 48000; omit to inherit |
Start and end (-ss, -to) | range: { start, end? } in seconds |
| 128 kbps mp3 | format: "mp3" (default is wav, pcm_s16le) |
What is the speech-to-text shape?
The docs call 16000 plus channels: "mono" the STT shape.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-stt-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "wav",
"channels": "mono",
"sample_rate": 16000
}'What about filters like volume or fades?
Not on this surface: detach accepts no filtergraph. For level and fades in a finished video use Timeline 1.0 fields such as audio.gain_db, output.fade_in_seconds and soundtrack.duck_db.
How does the job run?
Audio detach defaults to mode: "async". Pass mode: "sync" to wait up to 30 seconds for a 200 finished job, or you get 202 and poll GET /v1/jobs/:id/status and GET /v1/jobs/:id/result. There is no GET /v1/audio-detach/:id. Idempotency-Key is required. The result is kind: audio_detach with audio_url, duration_seconds, format, channels and sample_rate; sample_rate is null when you omitted it and the source rate was inherited. A warnings list may appear on the result.
Detach is not clip inspection. For probing, stills or transcripts use video inspect; for a new MP4 cut use video trim, described in the Audio detach docs.
Sources
Related posts
More in Media tools
- audio_parts_channel_mismatch: concat needs one channel layout
Timeline audio concat fails with audio_parts_channel_mismatch when parts mix channel layouts. Detach video audio as mono for every part.
- audio_parts_shorter_than_duration: fix a Timeline audio spine
Timeline 1.0 refuses audio.parts[] whose declared lengths sum to less than audio.duration_seconds. Add part length, lower the duration, or use silence mode.
- Combine more than 20 audio files: nest the concat jobs
Timeline audio concat takes 1 to 20 parts per job. For 45 files, concat in batches of 20 or fewer, then concat the batch outputs. Four jobs, $0.04 flat.
- Extract audio from a video URL: why example.com is refused
Audio detach only reads a video already on this workspace's media.sume.com. An outside URL fails unsupported_media_source, so import it first, then detach.
Written by Sume