ElevenLabs dubbing output: mono v1, stereo v2, and Sume channels

ElevenLabs dubbing v1 returns mono and v2 at most stereo. On Sume, detach keeps source channels or mono, and a concat of mixed layouts fails by design.

5 min readSume
All posts

ElevenLabs says its dubbing output is mono in v1 and at most stereo in Dubbing v2, and that multichannel sources are not preserved. On Sume, channel layout is something you choose at the detach step and then must keep consistent: detach offers source or mono, and a concat job fails with audio_parts_channel_mismatch if its parts disagree.

The ElevenLabs statement is from its Dubbing documentation (read 2026-10-10). Sume's rules are from Audio detach and Timeline audio. Neither page says anything about surround, so treat 5.1 sources as unsupported on both sides until you test.

Why channels matter in a dub

A dub replaces the speech track but usually keeps music and effects. If the original mix is stereo and your new voice track is mono, you have two decisions to make: whether to up-mix the voice, or to down-mix the original. Players cope with either, but broadcast and ad platforms often have strict channel rules, and a mismatch found at upload costs a re-render.

ElevenLabs makes the first decision for you: v1 is mono, v2 is at most stereo. Sume makes you declare it, which is more work but leaves no surprise.

Channel handling (ElevenLabs read 2026-10-10; Sume per docs.sume.com)
StepElevenLabs dubbingSume pipeline
Source with surroundMultichannel is not preservedDetach channels accepts source or mono; surround is not documented
Output layoutMono in v1; at most stereo in v2Whatever your TTS output and detach layouts are
Mixed layouts joinedNot applicable inside the dubConcat fails with audio_parts_channel_mismatch
Sample rateNot stated on that pageDetach: 16000, 44100 or 48000; TTS: 8000 to 48000

Detach settings for each purpose

Audio detach takes one Sume-hosted video and returns a new audio file for $0.01 per job. The default is sample-exact WAV (pcm_s16le); format: "mp3" gives 128 kbps. For speech recognition, the docs call 16000 Hz plus channels: "mono" the STT shape. For a track you will mix music under, keep channels: "source" so the stereo image survives.

A common pattern is two detaches of the same video: one 16 kHz mono file for transcription, one source-layout file for the music-and-effects bed. That is $0.02 in detach fees, and the two outputs never need to be joined.

Keeping the join consistent

When you join per-sentence TTS files with timeline audio concat, every part must share a layout. TTS output is controlled by output_format, which accepts containers mp3, wav and raw, sample rates from 8000 to 48000 Hz, and PCM encodings. Pick one wav setting once and use it for every sentence job, then the concat has no surprises.

If a failed join does happen, the error is named: audio_parts_channel_mismatch. That is a worker-side refusal, so it appears on the job result rather than at submit. Plan for a polling step before you assume the file exists.

A decision rule

Use the vendor's dub when you accept its layout rules and want one call. Use the Sume pipeline when you need to set layouts explicitly or to keep a stereo bed under a new voice. In the second case, write the layout decision into your job metadata so a later re-run reproduces the same settings.

Test one finished file in the real destination before batch work. Amazon audio ads, for example, have size and layout rules that a stereo WAV can break, as the mono WAV post explains.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume