Pull the voiceover out of a Reel MP4: audio detach, wav or mp3

To reuse the voice from a finished Reel, detach its audio for $0.01: wav by default, mp3 optional, a range up to 900 s, and a clean error for silent files.

5 min readSume
All posts

Sume's audio detach takes one video on media.sume.com and returns its audio track as a new file, for a flat $0.01 per job. The default is a sample-exact wav, mp3 at 128 kbps is the alternative, and the source video is left untouched. That makes it the first step when you want to reuse the voice of a finished Reel in a recut, a translation pass or a transcript.

The request

POST /v1/audio-detach requires video_url and an Idempotency-Key. The video must already be on your workspace's media host; import it first with POST /v1/media-imports. Optional fields are format (wav or mp3), range as { start, end? } in seconds, channels (source or mono) and sample_rate (16000, 44100 or 48000).

Audio detach options (read 2026-10-03)
FieldValuesUse it for
formatwav (default), mp3wav when the file will be joined or timed again
range{ start, end? }, up to 900 sOne section of a long Reel
channelssource (default), monoMono for speech tools
sample_rate16000, 44100, 4800016000 with mono is the speech-to-text shape

Limits that matter for Reels

The source can be up to 1800 seconds and the output up to 900 seconds, so a 3-minute Reel (Instagram's 2026 Reels guide, as reported by Metricool on 2026-10-03, lists lengths up to 3 minutes) is far inside both. A source without an audio track fails with detach_source_has_no_audio. Check first with video inspect and frames: false, which returns probe.has_audio without making stills.

A short script

The script submits the job in sync mode, which waits up to 30 seconds, and prints the job id and status; a finished job also carries the result, and an unfinished one is polled by that id.

import os, requests

r = requests.post(
    "https://api.sume.com/v1/audio-detach",
    headers={
        "Authorization": "Bearer " + os.environ["SUME_API_KEY"],
        "Idempotency-Key": "reel-voice-001",
    },
    json={
        "video_url": os.environ["REEL_URL"],
        "format": "wav",
        "mode": "sync",
    },
)
r.raise_for_status()
job = r.json()
print(job.get("request_id"), job.get("status"))

Where the file goes next

Use the wav as audio.url in Timeline 1.0, split it into beats with timeline audio, or send it to speech-to-text. Keep wav for anything that gets joined again; the docs note that mp3 re-adds priming padding at every edge when it is cut or joined.

What Sume does not do

Detach does not clean noise, separate voice from music or remove the original soundtrack from the picture. It copies out what the file contains. If the Reel has music baked under the voice, the detached file has both.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume