Extract audio from an AI avatar video as a WAV with one API call
Sume audio detach pulls the track of one hosted video into a WAV or MP3 for $0.01. The request, the 16 kHz mono option, the caps and the no-audio error.
To extract the audio from a finished avatar video on Sume, call POST /v1/audio-detach with the video's media.sume.com URL. You get a new audio file as a WAV by default (sample-exact PCM) or a 128 kbps MP3, and the video is untouched. It costs $0.01 per job, because it runs on the worker's media runtime with no model inference.
Use it when you want to transcribe the narration, reuse it under different footage, or check the voice track on its own.
The request
Idempotency-Key is required. The source must already be in your workspace on media.sume.com; there is no open-internet fetch, so import an outside file first with POST /v1/media-imports.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-intro-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "wav",
"sample_rate": 16000,
"channels": "mono"
}'Options and caps
The default mode is async; pass mode: "sync" to wait up to 30 seconds. Poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result for audio_url, duration_seconds and source_duration_seconds.
| Field | Values and limits |
|---|---|
format | wav (default) or mp3 at 128 kbps |
channels | source (default) or mono |
sample_rate | 16000, 44100 or 48000; omit to inherit the source |
range | { start, end? } in seconds; omit for the whole track |
| Source length | Up to 1,800 seconds |
| Output length | Up to 900 seconds; a longer track needs a range |
What it is good for
- Speech-to-text: 16 kHz mono WAV is the shape the docs name for transcription.
- Re-using a spoken take: detach once, then split it into pieces with timeline audio, which slices one file into up to 20 ranges.
- Quality checks: listen to the audio without the picture to find clipped words.
- Handing a clean track to a lip-sync step that takes still plus audio.
Reading the result
The result gives you audio_url on media.sume.com, the output duration_seconds, and the source_duration_seconds of the original. Compare the two when you passed a range: a mismatch tells you the cut was shorter than you expected. Download the file or pass the URL straight into another Sume route, since timeline audio requires Sume-hosted media. The job also appears in the jobs list, so a lost id can be recovered later.
Errors and limits
A source with no audio track fails with detach_source_has_no_audio. Check probe.has_audio first with video inspect and frames: false, which is unbilled for the probe. Off-host URLs are refused, and ffmpeg-style keys such as filtergraph, codec or crf are rejected with a 400, because detach does not accept raw encoder settings. If you need the audio of a video you did not make on Sume, you must have the right to use it before you import it.
Sources
Related posts
More in Media tools
- Extract audio from a video for transcription: 16 kHz mono wav
Detach the audio from a Sume-hosted video as 16 kHz mono wav, then send it to speech-to-text. Costs, caps and the errors to expect.
- Facebook Reels audio spec: AAC-LC, stereo, 48 kHz, 128 kbps+
Facebook Reels wants AAC Low Complexity, stereo, 48 kHz and 128 kbps or more. What a Sume probe shows, what trim keeps, and what you must verify yourself.
- FFmpeg drawtext on a hosted API: why Sume filters refuse it, and cues
Sume video-filter refuses drawtext and subtitles because they read files. To burn text onto a clip, send cues to video-captions: text, start and end seconds.
- Adjust video brightness, contrast, saturation by API: eq filter
Send eq=brightness=0.06:contrast=1.1:saturation=1.2 as the filtergraph of Sume video-filter. Ranges come from FFmpeg: brightness -1 to 1, saturation 0 to 3.
Written by Sume