YouTube lip sync dubbing: a pilot, and what you can do today

YouTube's lip sync dubbing is a pilot, not a setting you can turn on. What its help page says, and how Sume's still-plus-audio lip-sync routes differ from it.

4 min readSume
All posts

YouTube lip sync dubbing is described in YouTube's February 2026 post as a pilot only, so you cannot count on it for your own uploads. YouTube's help page defines it as an automatic dubbing feature that alters the speaker's lip movement to align with the dubbed audio. What Sume offers today is a different thing: a still image plus an audio clip in, a talking clip out.

What is the difference between the two?

They start from different inputs and change different things.

YouTube lip sync dubbing against Sume's lip-sync routes, read 2026-09-29.
YouTube lip sync dubbingSume lip-sync routes
StatusPilot only, per YouTube's postListed API routes
InputYour uploaded video with its speakerA still image or avatar plus Sume-hosted audio
What changesThe speaker's lip movement in your videoA new talking clip whose mouth follows the audio
Audio lengthNot stated in the sources read5 to 14.8 seconds on MiniMax H3 Max Lip Sync

Can Sume re-sync the lips of my existing video?

Not from what its docs list. Sume's models page says video models do not lip-sync to generated speech or to a later voice-over, and its lip-sync routes take a still, not a video. So a dubbed track laid under your original footage leaves the original mouth movement as it was. For a speaker on camera, the docs point to making new talking shots from a still plus speech instead.

Which Sume routes take a still and audio?

The models page lists VEED Fabric 1.0 at POST /v1/veed/fabric-1.0 and MiniMax H3 Max Lip Sync at POST /v1/minimax/h3-max/lip-sync. The second takes the same body as Fabric, with audio of 5 to 14.8 seconds, and the clip length follows the audio. Give exactly one visual source, image_url or an avatar, plus a Sume-hosted audio_url. Longer speech has to be cut into lines first; see lip sync for long video.

How do I choose between a re-synced video and a new talking clip?

Ask what the viewer needs to see. If the original speaker's face is on screen for most of the video, a dub with untouched lips will look off, and only the pilot feature addresses that, so it is outside your control today. If the face appears only in short moments, such as an intro or a presenter cutaway, remaking those moments as talking clips from a still and your dubbed speech is realistic, because each line is under the 14.8 second limit.

If the video is voiceover over screen recordings or product footage, there are no lips to match, and an ordinary dubbed audio track is the whole job.

What should I do while YouTube's version is a pilot?

Dub the audio and accept that lips will not match, or make talking-head shots from stills for the parts where a mouth is on screen. If you rely on YouTube's own auto dubbing, its help page lists the reasons a video may not be dubbed at all, and the Expressive Speech language list covers the eight languages named in YouTube's post.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume