Dubbed audio mono or stereo: keep the channel layout when you detach

ElevenLabs says dubbing v1 is mono and v2 up to stereo. On Sume, audio detach keeps source channels or mixes to mono, and concat needs matching layouts.

4 min readSume
All posts

Decide the channel layout before you dub, because mixing mono and stereo parts later causes errors and level jumps. ElevenLabs states that dubbing v1 output audio is mono and v2 is stereo at most. On Sume, audio detach has a channels option of source or mono, and timeline audio concat refuses parts whose channel layouts differ.

This post is about the handoff: you extract audio from a video, replace the speech somewhere, and bring it back without surprises.

What the vendor documents

The ElevenLabs dubbing doc says v2 covers 90+ languages with region-qualified dialects, handles up to 32 speakers per file, and accepts up to 3 GB per file through the API (1 GB and 180 minutes in the app; Dubbing Studio is 1 GB and 45 minutes). Output is a preview, most commonly 1080p, with mono audio for v1 and stereo at most for v2. Editing the transcript and regenerating through the API is Enterprise only.

ElevenLabs dubbing facts relevant to audio handoff (read 2026-10-03)
ItemValue
Languages (v2)90+, with regional dialects
Speakers per fileUp to 32
API file size3 GB per file
Audio channelsMono for v1, stereo at most for v2
Transcript edit and regenerate via APIEnterprise only

What Sume gives you

The audio detach docs let you extract a range from a video as wav or mp3 with channels set to source, which keeps the layout of the original, or mono, which mixes down. Choose a sample rate of 16000, 44100 or 48000. Source video up to 1800 seconds, output up to 900 seconds, $0.01 per job.

Joining is handled by timeline audio. Concat takes 1 to 20 parts and works in the sample domain without re-synthesis. If the parts have different channel layouts it returns audio_parts_channel_mismatch. That error is helpful: it tells you a mono dub was about to be glued to a stereo original.

  • Detach with channels source if you will lay the new voice over the original ambience.
  • Detach as mono when the target is speech-to-text or a mono voiceover.
  • Keep every concat part in the same layout.
  • Stay in wav until the final render; mp3 re-adds encoder padding.

A safe sequence

First detach the section you want to replace. Second, produce the new speech outside Sume, or with a Sume TTS job if you are writing the translation yourself; the sentence segments post shows how to keep timing. Third, import any off-host audio to Sume first, since timeline audio rejects off-host URLs. Fourth, check channel count and sample rate on every part, then concat or hand the audio url to a timeline render.

If the dub arrives in mono and you want stereo, decide where the stereo image should come from. A mono voice over a stereo music bed is normal, and the timeline soundtrack field handles that mix.

What Sume does not do

Sume does not run ElevenLabs dubbing, and there is no ElevenLabs model id in the catalog. This post only covers moving audio in and out around whatever dubbing tool you use.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume