Dubbed audio mono or stereo: keep the channel layout when you detach
ElevenLabs says dubbing v1 is mono and v2 up to stereo. On Sume, audio detach keeps source channels or mixes to mono, and concat needs matching layouts.

Decide the channel layout before you dub, because mixing mono and stereo parts later causes errors and level jumps. ElevenLabs states that dubbing v1 output audio is mono and v2 is stereo at most. On Sume, audio detach has a channels option of source or mono, and timeline audio concat refuses parts whose channel layouts differ.
This post is about the handoff: you extract audio from a video, replace the speech somewhere, and bring it back without surprises.
What the vendor documents
The ElevenLabs dubbing doc says v2 covers 90+ languages with region-qualified dialects, handles up to 32 speakers per file, and accepts up to 3 GB per file through the API (1 GB and 180 minutes in the app; Dubbing Studio is 1 GB and 45 minutes). Output is a preview, most commonly 1080p, with mono audio for v1 and stereo at most for v2. Editing the transcript and regenerating through the API is Enterprise only.
| Item | Value |
|---|---|
| Languages (v2) | 90+, with regional dialects |
| Speakers per file | Up to 32 |
| API file size | 3 GB per file |
| Audio channels | Mono for v1, stereo at most for v2 |
| Transcript edit and regenerate via API | Enterprise only |
What Sume gives you
The audio detach docs let you extract a range from a video as wav or mp3 with channels set to source, which keeps the layout of the original, or mono, which mixes down. Choose a sample rate of 16000, 44100 or 48000. Source video up to 1800 seconds, output up to 900 seconds, $0.01 per job.
Joining is handled by timeline audio. Concat takes 1 to 20 parts and works in the sample domain without re-synthesis. If the parts have different channel layouts it returns audio_parts_channel_mismatch. That error is helpful: it tells you a mono dub was about to be glued to a stereo original.
- Detach with channels source if you will lay the new voice over the original ambience.
- Detach as mono when the target is speech-to-text or a mono voiceover.
- Keep every concat part in the same layout.
- Stay in wav until the final render; mp3 re-adds encoder padding.
A safe sequence
First detach the section you want to replace. Second, produce the new speech outside Sume, or with a Sume TTS job if you are writing the translation yourself; the sentence segments post shows how to keep timing. Third, import any off-host audio to Sume first, since timeline audio rejects off-host URLs. Fourth, check channel count and sample rate on every part, then concat or hand the audio url to a timeline render.
If the dub arrives in mono and you want stereo, decide where the stereo image should come from. A mono voice over a stereo music bed is normal, and the timeline soundtrack field handles that mix.
What Sume does not do
Sume does not run ElevenLabs dubbing, and there is no ElevenLabs model id in the catalog. This post only covers moving audio in and out around whatever dubbing tool you use.
Sources
Related posts
More in Media tools
- Fake a second camera on one AI clip: a punch-in cut for Shorts
Crop one AI clip with video-filter, then cut between wide and tight versions in Timeline 1.0 at the same moments, so a Short keeps moving without regenerating.
- FFmpeg 9.0.2 Lei: what is new, and what hosted trim exposes
FFmpeg 9.0.2 shipped 2026-09-18. Sume runs ffmpeg for you but accepts no ffmpeg flags: trim takes start, duration, precision, audio and output.
- Final Cut Pro 12.4 Cinematic edits: export flat, then caption
Final Cut Pro 12.4 edits Cinematic focus after the fact. Sume reads a public MP4 URL, so export a flattened file first, then trim or caption it.
- Firefly Video Editor Quick Cut vs a scripted first cut
Adobe's Firefly Video Editor has Quick Cut for a first cut from raw footage. Sume can inspect clips, trim, and join them on a timeline, but picks no shots.
Written by Sume