Dub a video, keep the music: Sume has no stem splitter, so do this

Sume can detach a video's audio and re-voice a script, but its docs list no vocal and music separation. A clean workaround with a new music bed and ducking.

6 min readSume
All posts

Sume cannot split a finished mix into a voice stem and a music stem, so if your source video has music baked under the speech, a dub will not keep that exact music. What you can do on Sume is replace the whole audio bed: detach the audio, transcribe and translate the speech yourself, generate the new voice, create a fresh music bed with the Music Router, and let Timeline 1.0 lay the voice over the music with ducking. If you have the original stems from the editor, use them and skip all of this.

This post walks through both cases and is explicit about what each Sume step does and does not do.

What the Sume docs list

None of these is a separator. Audio detach returns the whole track. There is no Sume endpoint in those docs that removes speech from music or music from speech, so say so to your client rather than promise it.

From Sume's Audio detach, Timeline audio, Music Router and Timeline 1.0 pages, read 2026-10-03.
StepSurfaceWhat it doesLimit
Pull the trackPOST /v1/audio-detachExtracts the video's audio as wav or mp3; video untouchedSource up to 1,800 s; output up to 900 s; $0.01 per job
Join or slice audioPOST /v1/timeline-1.0/audioConcat up to 20 parts or split up to 20 ranges, sample-domain$0.01 flat per job
New music bedPOST /v1/music-router/generatePrompt in, audio out; sume/music-auto picks Lyria 3.5 todayNo duration field; steer length in the prompt
Lay voice over bedPOST /v1/timeline-1.0/rendersoundtrack with gain_db, loop, fade_out_seconds up to 10, duck_db 0 to 20Needs a real audio spine

Case 1: you have the stems

If the editor can export a voice-only track and a music-only track, you are in the easy case. Translate the voice stem's transcript, generate the new voice, join takes if you made several with timeline audio, and use the original music stem as the soundtrack in the render. Keep the music's gain_db and duck_db close to the original mix. Sume's Timeline render takes the music as a hosted url, so import it first, as every Sume media call requires.

Case 2: baked-in music

If you only have the final mix, the honest options are three. Ask for the stems. Dub only the speech and accept a new bed, which is a creative change you should show the client before you do it. Or leave music-heavy sections alone and dub only the talking parts, cutting between original and dubbed with timeline audio splits.

For the new bed, write the music brief in positive terms, because the Music Router does not support a non-empty negative_prompt, and name the length in the prompt, because it rejects duration and duration_seconds. Something like: warm lo-fi hip hop, 84 BPM, instrumental, no vocals, a 60-second track. The result is an audio artifact on the job's result.artifacts[].

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: dub-bed-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Warm lo-fi hip hop, 84 BPM, C minor. Dusty Rhodes chords, brushed drums, subtle bass. A 60-second track. Instrumental, no vocals, nothing that competes with a spoken voiceover."
  }'

What to tell the client

Put the limits in writing before you start. A dub that replaces the audio bed is a different deliverable from a dub that preserves the original music; clients often assume the second. Say which one you are making, show a 15-second sample with the new bed, and get approval before you spend on the full run.

The cost side is small: detach is $0.01, a bed is one music generation at the fixed Music price, the voice is $47.50 per 1M characters, and a join is $0.01. The time is in the translation and the review, which no API removes.

Laying the dubbed voice over the bed

In Timeline 1.0 the voice is the audio spine (audio.url or audio.parts[]) and the bed is soundtrack. A duck_db between 0 and 20 lowers the bed under speech, fade_out_seconds up to 10 lets it end cleanly, and loop repeats a short bed under a longer voice. Ducking needs a real spine, so it does not work on a silent one. The spine length is audio.duration_seconds, up to 1,800.

Two cautions. A generated bed is not the client's original music, so the brand may need to approve it. And lip-synced footage will not match a new language: Sume's models overview says video models do not lip-sync to generated TTS or to a later voice-over, so a dub like this suits voice-over-style videos, b-roll and screen recordings better than a talking head.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume