Extract audio from MP4 to WAV or MP3: the detach defaults
Sume audio detach returns WAV (pcm_s16le, sample-exact) by default or MP3 at 128 kbps. Pick WAV for speech-to-text and timeline use, MP3 when file size matters.

Audio detach gives you WAV by default: pcm_s16le, sample-exact. Set format: "mp3" for a 128 kbps MP3 instead. WAV is the shape the docs say timeline_create audio.url, POST /v1/timeline-1.0/audio and speech-to-text want; MP3 is the smaller file.
Details are from the audio detach docs, read 2026-10-01. HeyGen's Speech Cleanup post, read the same day, says it accepts MP4, MOV or WEBM uploads; Sume's detach takes a video already on media.sume.com, so import first with POST /v1/media-imports.
What are the format and audio options?
All fields are optional except video_url.
| Field | Values | Default |
|---|---|---|
format | wav (pcm_s16le), mp3 (128 kbps) | wav |
channels | source, mono | source |
sample_rate | 16000, 44100, 48000 | Inherit the source |
range | { start, end? } seconds | Whole track |
How big is each format?
Arithmetic on the stated encodings, assuming a stereo 48 kHz source: 16-bit PCM is 48,000 samples x 2 channels x 2 bytes, about 192 KB per second or 11.5 MB per minute. MP3 at 128 kbps is 16 KB per second, about 0.96 MB per minute. At the 900-second output cap that is roughly 173 MB for WAV and 14 MB for MP3. These are calculations, not measured files; your source's channels and rate change the WAV figure.
Which one should I pick?
Choose WAV when the audio goes on to speech-to-text or a Timeline spine; 16000 with channels: "mono" is the speech-to-text shape the docs name. Choose MP3 for a listening copy. If the request returns unsupported_media_type, the HEAD of the source was not a video; see also extract audio from a video URL and WAV or MP3 for general trade-offs.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-mp3-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "mp3"
}'Sources
Related posts
More in Developers
- Extract audio segment start beyond video length: error and fix
Audio detach fails detach_start_past_source when range.start is past the probed duration. It differs from audio_detach_range_empty. Check duration first.
- Devin Desktop ACP session/new mcp_servers with Sume's remote URL
Devin Desktop 3.10.35 uses MCP servers an ACP client passes in session/new. Pass Sume's streamable HTTP URL and an API-key header, not editor config.
- ElevenLabs API timeout: cascade_timeout_seconds vs Sume waits
ElevenLabs added cascade_timeout_seconds (2-15 s) for Speech Engine retries. Sume's wait_timeout_seconds is a different knob: a 0-30 s HTTP wait on a job.
- ElevenLabs CLI 1.0: scripting audio with Sume over plain HTTP
ElevenLabs CLI v1.0.0 turns every API operation into a subcommand. For Sume audio, script curl against /v1/jobs and join clips for $0.01 flat per job.
Written by Sume