YouTube preferred language setting: keep the source audio clean
YouTube now lets viewers set a preferred language for dubs. If you add your own dub, start from a clean speech track and detach it from the video first.

YouTube's blog describes a Preferred Language setting that lets a multilingual viewer choose how they hear content, with the default based on their watch history. It also says creators keep full control to provide custom dubs or turn the feature off. So a viewer's choice can pick between your original track and a dub, and a dub you supply is yours to prepare.
Why the source track matters
A dub starts from the original speech. If music or sound effects are mixed into the track, a transcript of it is noisier and a replacement voice has to compete with whatever you kept. The most useful preparation is to keep a clean dialogue take, and to make a separate version without the bed.
Detach once
Sume can pull the sound out of a video that sits in your workspace. POST /v1/audio-detach returns a new audio file and leaves the video unchanged. Use format: "wav" for a sample-exact file, and channels: "mono" with sample_rate: 16000 if the next step is speech-to-text. A video with no audio track fails with detach_source_has_no_audio, so check probe.has_audio with video inspect first.
Then build the dub
From the detached file you transcribe, translate and speak the new track, as in a normal dub. Keep each language as its own audio file, named for its language, so you upload the right one for each track. Detach once and split with timeline audio when you need many ranges, since one detach is cheaper than many.
What the vendor says versus what you do, read 2026-10-06:
| Point | YouTube blog | Your action on Sume |
|---|---|---|
| Viewer control | Preferred Language setting | None, it is a viewer setting |
| Creator control | Custom dubs or turn off dubbing | Make the dub files yourself |
| Source audio | Not stated | Detach a clean wav from the video |
| Limits | Not stated | Source 1800 s, output 900 s per detach |
A clean-track checklist
Record dialogue without music if you can, and keep the music as a separate file. Name every take. Keep a text script that matches what is said, since it is the starting point for translation.
If the only copy of the video has music under the voice, a transcript still works, but a replacement voice will have no music unless you add a bed. Plan that bed as a separate step, and decide who owns the licence.
Sume does not upload to YouTube and does not read your channel settings. You publish the audio yourself. Check YouTube's own help pages for how to add audio tracks, since this page only covers the preparation.
Sources
Related posts
More in Use cases
- YouTube Shorts series covers: one template, one edit per episode
YouTube is rolling out Shorts series with covers. Make one template with Ideogram 4.5, then change only the episode number with one edit call per episode.
- Shorts series: a season of 30-second episodes with Seedance 2.5
Shorts series are rolling out. A season of eight 30-second episodes costs $138.72 on Seedance 2.5 at 720p and $30.00 on Wan 3.0. The math and a batch loop.
- YouTube show episode numbers follow publish date: render in that order
YouTube assigns Shorts episode numbers by publish date unless you set a manual playlist order. Plan your render queue so numbers match your story.
- Will a YouTube show accept audio-led Shorts with no real visuals?
YouTube says videos with no visuals likely will not be considered a show. How to give narrated, audio-led Shorts a real picture track before you build a series.
Written by Sume