Dub a video, keep the music: Sume has no stem splitter, so do this
Sume can detach a video's audio and re-voice a script, but its docs list no vocal and music separation. A clean workaround with a new music bed and ducking.

Sume cannot split a finished mix into a voice stem and a music stem, so if your source video has music baked under the speech, a dub will not keep that exact music. What you can do on Sume is replace the whole audio bed: detach the audio, transcribe and translate the speech yourself, generate the new voice, create a fresh music bed with the Music Router, and let Timeline 1.0 lay the voice over the music with ducking. If you have the original stems from the editor, use them and skip all of this.
This post walks through both cases and is explicit about what each Sume step does and does not do.
What the Sume docs list
None of these is a separator. Audio detach returns the whole track. There is no Sume endpoint in those docs that removes speech from music or music from speech, so say so to your client rather than promise it.
| Step | Surface | What it does | Limit |
|---|---|---|---|
| Pull the track | POST /v1/audio-detach | Extracts the video's audio as wav or mp3; video untouched | Source up to 1,800 s; output up to 900 s; $0.01 per job |
| Join or slice audio | POST /v1/timeline-1.0/audio | Concat up to 20 parts or split up to 20 ranges, sample-domain | $0.01 flat per job |
| New music bed | POST /v1/music-router/generate | Prompt in, audio out; sume/music-auto picks Lyria 3.5 today | No duration field; steer length in the prompt |
| Lay voice over bed | POST /v1/timeline-1.0/render | soundtrack with gain_db, loop, fade_out_seconds up to 10, duck_db 0 to 20 | Needs a real audio spine |
Case 1: you have the stems
If the editor can export a voice-only track and a music-only track, you are in the easy case. Translate the voice stem's transcript, generate the new voice, join takes if you made several with timeline audio, and use the original music stem as the soundtrack in the render. Keep the music's gain_db and duck_db close to the original mix. Sume's Timeline render takes the music as a hosted url, so import it first, as every Sume media call requires.
Case 2: baked-in music
If you only have the final mix, the honest options are three. Ask for the stems. Dub only the speech and accept a new bed, which is a creative change you should show the client before you do it. Or leave music-heavy sections alone and dub only the talking parts, cutting between original and dubbed with timeline audio splits.
For the new bed, write the music brief in positive terms, because the Music Router does not support a non-empty negative_prompt, and name the length in the prompt, because it rejects duration and duration_seconds. Something like: warm lo-fi hip hop, 84 BPM, instrumental, no vocals, a 60-second track. The result is an audio artifact on the job's result.artifacts[].
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dub-bed-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Warm lo-fi hip hop, 84 BPM, C minor. Dusty Rhodes chords, brushed drums, subtle bass. A 60-second track. Instrumental, no vocals, nothing that competes with a spoken voiceover."
}'
What to tell the client
Put the limits in writing before you start. A dub that replaces the audio bed is a different deliverable from a dub that preserves the original music; clients often assume the second. Say which one you are making, show a 15-second sample with the new bed, and get approval before you spend on the full run.
The cost side is small: detach is $0.01, a bed is one music generation at the fixed Music price, the voice is $47.50 per 1M characters, and a join is $0.01. The time is in the translation and the review, which no API removes.
Laying the dubbed voice over the bed
In Timeline 1.0 the voice is the audio spine (audio.url or audio.parts[]) and the bed is soundtrack. A duck_db between 0 and 20 lowers the bed under speech, fade_out_seconds up to 10 lets it end cleanly, and loop repeats a short bed under a longer voice. Ducking needs a real spine, so it does not work on a silent one. The spine length is audio.duration_seconds, up to 1,800.
Two cautions. A generated bed is not the client's original music, so the brand may need to approve it. And lip-synced footage will not match a new language: Sume's models overview says video models do not lip-sync to generated TTS or to a later voice-over, so a dub like this suits voice-over-style videos, b-roll and screen recordings better than a talking head.
Sources
Related posts
More in Use cases
- Dub a video where two languages are spoken: keep the original lines
Sume's speech-to-text takes one language hint per call. For a mixed-language video, transcribe, translate only some lines, and splice the rest back in.
- DV360 video creative specs: H.264, 20 Mbps, -24 LKFS audio
Display & Video 360 asks for H.264 at 20 Mbps or more, 23.98 or 29.97 fps, 48 kHz audio at -24 LKFS. Which of those Sume's trim and timeline can and cannot set.
- eBay PictureURL: first URL is the Gallery image, 3,975-char cap
eBay allows up to 24 picture URLs, uses the first as the Gallery image, and caps all PictureURL values at 3,975 characters. How to order and size a Sume batch.
- eBay Promoted Listings video ads: review time and a sound-off clip
eBay reviews Promoted Listings video before it runs: most in 48 hours, up to 7 business days. Aim for 5 to 15 seconds, sound off, and build the clip with Sume.
Written by Sume