YouTube lip sync for auto dubs: what to prepare before you get access
YouTube is piloting lip sync on auto dubs. Prepare clean speech audio and a 5 to 14.8 second lip-synced clip workflow on Sume while you wait.

YouTube's own blog says it is testing a Lip Sync pilot that subtly matches the speaker's lip movements to the translated audio. The post does not give a start date for everyone, so a creator who wants a dubbed video that looks right today cannot wait on it. What you can do now is prepare the two inputs any lip-sync step needs: a clean dubbed voice track, and a short clip of the speaker.
What Sume gives you today
Sume does not offer a YouTube feature, and its lip-sync surface is a different product. The MiniMax H3 Max lip-sync route takes one still (or a ready avatar) plus a Sume-hosted audio file of 5 to 14.8 seconds and at most 10 MB, and returns a talking clip. VEED Fabric 1.0 takes the same still-plus-audio body. The API refuses a duration_seconds outside 5 to 14.8 and never clamps it, because the provider silently clips audio after 14.8 seconds.
Prepare the inputs
Treat it as a cut-down version of the dub. Your dub is one TTS job per sentence in the target language, so every line is already a short file. Join lines that fit the window and send each to lip sync, then place the clips on a timeline. A full-length dubbed video is a different route, and it needs timeline rather than a single lip-sync job.
- Write the translated script yourself, one sentence per line, so each TTS file is easy to time.
- Set
languageon each TTS request, and keep the voice's language matching it. - Check each clip's length before you submit it, since lip sync accepts 5 to 14.8 seconds.
- Keep the original audio on the video until the dubbed take passes your listen test.
Facts from the two sources, read 2026-10-06:
| Item | YouTube (vendor blog) | Sume (docs) |
|---|---|---|
| Status | Lip Sync pilot being tested | Public routes, no pilot |
| Input | Your uploaded video | One still or avatar plus Sume-hosted audio |
| Audio length | Not stated | 5 to 14.8 s, at most 10 MB |
| Output | Dubbed video, lips adjusted | One talking clip per job |
What not to promise
Do not tell a client that a Sume clip is the same as YouTube's pilot. The blog gives one sentence about the pilot and no list of channels, languages or lengths, so anything more is a guess. Say that you prepared clean dubbed audio and a lip-synced test clip, and that YouTube's own feature is separate.
Also keep cost in mind. Lip sync is billed from the audio length, so run it only on the shots where a mouth shows. A wide shot, a cutaway or a screen recording needs the dubbed voice and nothing else.
If you only need the dubbed voice, YouTube's own auto dub or your own track is the lower-effort path. Use lip sync only for the shots where a mouth is visible in frame.
Sources
Related posts
More in Use cases
- YouTube preferred language setting: keep the source audio clean
YouTube now lets viewers set a preferred language for dubs. If you add your own dub, start from a clean speech track and detach it from the video first.
- YouTube Shorts series covers: one template, one edit per episode
YouTube is rolling out Shorts series with covers. Make one template with Ideogram 4.5, then change only the episode number with one edit call per episode.
- Shorts series: a season of 30-second episodes with Seedance 2.5
Shorts series are rolling out. A season of eight 30-second episodes costs $138.72 on Seedance 2.5 at 720p and $30.00 on Wan 3.0. The math and a batch loop.
- YouTube show episode numbers follow publish date: render in that order
YouTube assigns Shorts episode numbers by publish date unless you set a manual playlist order. Plan your render queue so numbers match your story.
Written by Sume