YouTube preferred language setting: keep the source audio clean

YouTube now lets viewers set a preferred language for dubs. If you add your own dub, start from a clean speech track and detach it from the video first.

4 min readSume
All posts

YouTube's blog describes a Preferred Language setting that lets a multilingual viewer choose how they hear content, with the default based on their watch history. It also says creators keep full control to provide custom dubs or turn the feature off. So a viewer's choice can pick between your original track and a dub, and a dub you supply is yours to prepare.

Why the source track matters

A dub starts from the original speech. If music or sound effects are mixed into the track, a transcript of it is noisier and a replacement voice has to compete with whatever you kept. The most useful preparation is to keep a clean dialogue take, and to make a separate version without the bed.

Detach once

Sume can pull the sound out of a video that sits in your workspace. POST /v1/audio-detach returns a new audio file and leaves the video unchanged. Use format: "wav" for a sample-exact file, and channels: "mono" with sample_rate: 16000 if the next step is speech-to-text. A video with no audio track fails with detach_source_has_no_audio, so check probe.has_audio with video inspect first.

Then build the dub

From the detached file you transcribe, translate and speak the new track, as in a normal dub. Keep each language as its own audio file, named for its language, so you upload the right one for each track. Detach once and split with timeline audio when you need many ranges, since one detach is cheaper than many.

What the vendor says versus what you do, read 2026-10-06:

Preferred language and your own dub, read 2026-10-06
PointYouTube blogYour action on Sume
Viewer controlPreferred Language settingNone, it is a viewer setting
Creator controlCustom dubs or turn off dubbingMake the dub files yourself
Source audioNot statedDetach a clean wav from the video
LimitsNot statedSource 1800 s, output 900 s per detach

A clean-track checklist

Record dialogue without music if you can, and keep the music as a separate file. Name every take. Keep a text script that matches what is said, since it is the starting point for translation.

If the only copy of the video has music under the voice, a transcript still works, but a replacement voice will have no music unless you add a bed. Plan that bed as a separate step, and decide who owns the licence.

Sume does not upload to YouTube and does not read your channel settings. You publish the audio yourself. Check YouTube's own help pages for how to add audio tracks, since this page only covers the preparation.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume