Extract audio from an AI avatar video as a WAV with one API call

Sume audio detach pulls the track of one hosted video into a WAV or MP3 for $0.01. The request, the 16 kHz mono option, the caps and the no-audio error.

4 min readSume
All posts

To extract the audio from a finished avatar video on Sume, call POST /v1/audio-detach with the video's media.sume.com URL. You get a new audio file as a WAV by default (sample-exact PCM) or a 128 kbps MP3, and the video is untouched. It costs $0.01 per job, because it runs on the worker's media runtime with no model inference.

Use it when you want to transcribe the narration, reuse it under different footage, or check the voice track on its own.

The request

Idempotency-Key is required. The source must already be in your workspace on media.sume.com; there is no open-internet fetch, so import an outside file first with POST /v1/media-imports.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: detach-intro-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "format": "wav",
    "sample_rate": 16000,
    "channels": "mono"
  }'

Options and caps

The default mode is async; pass mode: "sync" to wait up to 30 seconds. Poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result for audio_url, duration_seconds and source_duration_seconds.

From Audio detach, read 2026-10-01.
FieldValues and limits
formatwav (default) or mp3 at 128 kbps
channelssource (default) or mono
sample_rate16000, 44100 or 48000; omit to inherit the source
range{ start, end? } in seconds; omit for the whole track
Source lengthUp to 1,800 seconds
Output lengthUp to 900 seconds; a longer track needs a range

What it is good for

  • Speech-to-text: 16 kHz mono WAV is the shape the docs name for transcription.
  • Re-using a spoken take: detach once, then split it into pieces with timeline audio, which slices one file into up to 20 ranges.
  • Quality checks: listen to the audio without the picture to find clipped words.
  • Handing a clean track to a lip-sync step that takes still plus audio.

Reading the result

The result gives you audio_url on media.sume.com, the output duration_seconds, and the source_duration_seconds of the original. Compare the two when you passed a range: a mismatch tells you the cut was shorter than you expected. Download the file or pass the URL straight into another Sume route, since timeline audio requires Sume-hosted media. The job also appears in the jobs list, so a lost id can be recovered later.

Errors and limits

A source with no audio track fails with detach_source_has_no_audio. Check probe.has_audio first with video inspect and frames: false, which is unbilled for the probe. Off-host URLs are refused, and ffmpeg-style keys such as filtergraph, codec or crf are rejected with a 400, because detach does not accept raw encoder settings. If you need the audio of a video you did not make on Sume, you must have the right to use it before you import it.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume