YouTube multi-language audio: add dubbed tracks, made with AI

YouTube multi-language audio lets one video carry dubbed tracks you upload yourself. The rules, the upload steps, and how to make each track with AI.

5 min readSume
All posts

YouTube's multi-language audio lets you add audio tracks in other languages to a single video or Short, and each viewer hears the track for their preferred language by default. You upload your own dubbed, audio-only tracks in YouTube Studio: the feature doesn't create them, and it's available to creators with access to Advanced features.

The YouTube rules below come from its help page Add Multi-language features to your videos, read on 2026-09-28. The steps for making a track with AI use the speech schemas in the Sume API reference and Sume's Timeline audio docs.

How does YouTube multi-language audio work?

One upload, several soundtracks. The rules that matter before you make any audio:

From YouTube Help, Add Multi-language features to your videos, read 2026-09-28.
RuleWhat YouTube's page says
What it addsAudio tracks in different languages on a single video or Short
What viewers hearThe track for their preferred language by default; they can switch in the player settings
Who makes the tracksYou do: it does not create them, unlike automatic dubbing
Existing auto dubDelete the auto dub in that language before uploading your own
EligibilityCreators with access to Advanced features
FileA supported audio-only file format, roughly the same length as the video

How do I add a multi-language audio track on YouTube?

YouTube's steps, in YouTube Studio on a computer:

  • From the left menu, select Languages, then click the video.
  • Click Add Language and pick the language.
  • Next to Dub, click Add, then Select file, and choose the audio file.
  • Click Publish. To replace a track later, click Delete under Audio for that language and upload again.

How do I make a dubbed audio track with AI?

Four steps per language: get the original words as text, translate them, speak the translation, and check the length.

  • Text: use your script, or transcribe the original with POST /v1/stt-1.0/transcribe, which takes a public HTTPS audio_url. For a video already on Sume, POST /v1/audio-detach pulls its audio track out as wav or mp3 first.
  • Translation: do it yourself or with an agent, and review it; YouTube recommends second-language tracks match the original as closely as possible.
  • Speech: send the translated transcript to POST /v1/tts-1.0/generate with a voice and the track's language code in language (BCP-47 / ISO-639); omitted, it defaults to English. Multilingual text to speech API covers which languages to test and the voice-language check.
  • Length: YouTube wants roughly the video's length. If the track runs long or short, set generation_config.speed between 0.6 and 1.5 and render again. For a script too long for one request, speak it in parts and join them with Timeline audio.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: yt-dub-es-001" \
  -d '{
    "transcript": "Hoy te enseño a configurar el panel en un minuto.",
    "avatar_handle": "acme",
    "language": "es",
    "generation_config": { "speed": 1.1 }
  }'

What does a dubbed track cost, and what are the limits?

Each step bills on its own, plus a 5.5% agent fee by default. TTS returns mp3 by default, or wav; confirm that the file type is one YouTube Studio accepts before you upload.

From the STT and TTS schemas in the Sume API reference, Timeline audio and API pricing, read 2026-09-28.
StepLimitPrice
Transcribe the original (STT 1.0)Public HTTPS audio_url; up to 10 minutes per request$0.01 per audio minute
Speak the translation (TTS 1.0)Up to 20,000 characters; audio over 1,200 s fails; speed 0.6–1.5$0.0475 per 1,000 characters
Join parts (Timeline audio)1–20 Sume-hosted parts, joined with no silence at the seamsSee API pricing

Is this the same as dubbing the video itself?

No. A multi-language audio track only swaps the sound; the picture, including the speaker's lips, stays the original. To re-render the video over a new voice, see Translate a video's voiceover, and Lip sync vs dubbing explains when matching lips matters.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume