YouTube lip sync dubbing: a pilot, and what you can do today
YouTube's lip sync dubbing is a pilot, not a setting you can turn on. What its help page says, and how Sume's still-plus-audio lip-sync routes differ from it.
YouTube lip sync dubbing is described in YouTube's February 2026 post as a pilot only, so you cannot count on it for your own uploads. YouTube's help page defines it as an automatic dubbing feature that alters the speaker's lip movement to align with the dubbed audio. What Sume offers today is a different thing: a still image plus an audio clip in, a talking clip out.
What is the difference between the two?
They start from different inputs and change different things.
| YouTube lip sync dubbing | Sume lip-sync routes | |
|---|---|---|
| Status | Pilot only, per YouTube's post | Listed API routes |
| Input | Your uploaded video with its speaker | A still image or avatar plus Sume-hosted audio |
| What changes | The speaker's lip movement in your video | A new talking clip whose mouth follows the audio |
| Audio length | Not stated in the sources read | 5 to 14.8 seconds on MiniMax H3 Max Lip Sync |
Can Sume re-sync the lips of my existing video?
Not from what its docs list. Sume's models page says video models do not lip-sync to generated speech or to a later voice-over, and its lip-sync routes take a still, not a video. So a dubbed track laid under your original footage leaves the original mouth movement as it was. For a speaker on camera, the docs point to making new talking shots from a still plus speech instead.
Which Sume routes take a still and audio?
The models page lists VEED Fabric 1.0 at POST /v1/veed/fabric-1.0 and MiniMax H3 Max Lip Sync at POST /v1/minimax/h3-max/lip-sync. The second takes the same body as Fabric, with audio of 5 to 14.8 seconds, and the clip length follows the audio. Give exactly one visual source, image_url or an avatar, plus a Sume-hosted audio_url. Longer speech has to be cut into lines first; see lip sync for long video.
How do I choose between a re-synced video and a new talking clip?
Ask what the viewer needs to see. If the original speaker's face is on screen for most of the video, a dub with untouched lips will look off, and only the pilot feature addresses that, so it is outside your control today. If the face appears only in short moments, such as an intro or a presenter cutaway, remaking those moments as talking clips from a still and your dubbed speech is realistic, because each line is under the 14.8 second limit.
If the video is voiceover over screen recordings or product footage, there are no lips to match, and an ordinary dubbed audio track is the whole job.
What should I do while YouTube's version is a pilot?
Dub the audio and accept that lips will not match, or make talking-head shots from stills for the parts where a mouth is on screen. If you rely on YouTube's own auto dubbing, its help page lists the reasons a video may not be dubbed at all, and the Expressive Speech language list covers the eight languages named in YouTube's post.
Sources
Related posts
More in Use cases
- YouTube real-time auto dubbing for live streams: what we know
YouTube says real-time auto dubbing is coming for live. Only the finished-video dubbing rules are published; here is the gap, and what Sume can and cannot do.
- YouTube Shopping affiliate in 35 countries: video ideas to batch
YouTube says its shopping affiliate program reaches 35 countries by year-end. Five product video formats to batch, with the Sume length and shape limits.
- How to make a YouTube Shorts series: seasons and episodes
A YouTube Shorts series is a playlist of Shorts only, set up as a show with seasons and episodes. Rules from YouTube's page, and how to render episodes.
- Does upscaling video with AI need disclosure on YouTube?
No, YouTube lists video sharpening, upscaling and repair as not needing disclosure. The exception is an edit that also changes what a real event shows.
Written by Sume