Pull the voiceover out of a Reel MP4: audio detach, wav or mp3
To reuse the voice from a finished Reel, detach its audio for $0.01: wav by default, mp3 optional, a range up to 900 s, and a clean error for silent files.

Sume's audio detach takes one video on media.sume.com and returns its audio track as a new file, for a flat $0.01 per job. The default is a sample-exact wav, mp3 at 128 kbps is the alternative, and the source video is left untouched. That makes it the first step when you want to reuse the voice of a finished Reel in a recut, a translation pass or a transcript.
The request
POST /v1/audio-detach requires video_url and an Idempotency-Key. The video must already be on your workspace's media host; import it first with POST /v1/media-imports. Optional fields are format (wav or mp3), range as { start, end? } in seconds, channels (source or mono) and sample_rate (16000, 44100 or 48000).
| Field | Values | Use it for |
|---|---|---|
| format | wav (default), mp3 | wav when the file will be joined or timed again |
| range | { start, end? }, up to 900 s | One section of a long Reel |
| channels | source (default), mono | Mono for speech tools |
| sample_rate | 16000, 44100, 48000 | 16000 with mono is the speech-to-text shape |
Limits that matter for Reels
The source can be up to 1800 seconds and the output up to 900 seconds, so a 3-minute Reel (Instagram's 2026 Reels guide, as reported by Metricool on 2026-10-03, lists lengths up to 3 minutes) is far inside both. A source without an audio track fails with detach_source_has_no_audio. Check first with video inspect and frames: false, which returns probe.has_audio without making stills.
A short script
The script submits the job in sync mode, which waits up to 30 seconds, and prints the job id and status; a finished job also carries the result, and an unfinished one is polled by that id.
import os, requests
r = requests.post(
"https://api.sume.com/v1/audio-detach",
headers={
"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "reel-voice-001",
},
json={
"video_url": os.environ["REEL_URL"],
"format": "wav",
"mode": "sync",
},
)
r.raise_for_status()
job = r.json()
print(job.get("request_id"), job.get("status"))Where the file goes next
Use the wav as audio.url in Timeline 1.0, split it into beats with timeline audio, or send it to speech-to-text. Keep wav for anything that gets joined again; the docs note that mp3 re-adds priming padding at every edge when it is cut or joined.
What Sume does not do
Detach does not clean noise, separate voice from music or remove the original soundtrack from the picture. It copies out what the file contains. If the Reel has music baked under the voice, the detached file has both.
Sources
Related posts
More in Media tools
- Reel music bed: prompt section markers for a 3-minute Reel
Music Router rejects a duration field, so steer length in the prompt. Write [0:00-0:30] section markers for a 3-minute Reel bed and loop it in Timeline.
- Reel voiceover too long for 60 seconds? Speed setting and re-render
A 68-second TTS read against a 60-second Reel: Sume accepts generation_config.speed from 0.6 to 1.5. The arithmetic, and when a rewrite is better.
- Clipdrop remove-background API: 60 requests a minute vs Sume RMBG
Clipdrop's remove-background API takes a 30 MB, 25 MP upload at 60 requests a minute per key. Sume RMBG takes a public HTTPS image_url and runs as a job.
- Resolve 21.1 multicam up to 25 angles vs Sume timeline slots
Resolve 21.1 adds multicam viewing with shortcuts for up to 25 angles. Sume has no multicam; Timeline 1.0 stacks 1 to 200 clips on one audio spine.
Written by Sume