Difference between lip sync and dubbing: which do you need?
Dubbing replaces the speech in a video; lip sync matches a mouth to the audio. Lip-sync dubbing does both. What each means, and which one you need.

Dubbing replaces what a video says: the original speech is swapped for a new recording, for example in another language, while the picture stays as it was. Lip sync is the match between a speaker's mouth and the sound, so the lips move with the words. Lip-sync dubbing does both: the new speech and the mouth line up, so the speaker looks as if they were saying the new words.
Sume handles the two separately: TTS 1.0 speaks new lines, and VEED Fabric 1.0 lip-syncs a still image or an avatar to that audio. Fabric takes no video input, so Sume can't re-time the mouth in footage you already have. The Sume facts come from the Models overview and Timeline 1.0 docs and the Sume API reference, read on 2026-09-28; the definitions are general. Anything called current behavior is read from Sume's code.
What is dubbing, and how is it different from a voice-over?
Dubbing puts new speech in the speaker's place: the viewer is meant to hear the new voice as the person on screen. A voice-over adds a narrator who is not the person on screen, and it can play over the original sound turned down. Both change only the audio. Neither moves anyone's lips, so when a speaking face is on screen, a dub or a voice-over leaves the mouth making the shapes of the old words.
What is lip sync?
Lip sync is making mouth movements and sound match. A performer lip-syncs by miming to a recording, a dubbing team times new speech to the lips already on screen, and AI lip sync generates the mouth movement from the audio itself, so the speech has to exist first.
On Sume, lip sync is VEED Fabric 1.0: a still image or a ready avatar plus Sume-hosted audio in, a talking clip out. A video model with narration laid underneath is not lip sync: Sume's model guide says video models do not lip-sync to generated TTS or to a later voice-over. Lip sync API: turn an image and audio into a talking clip covers the request and its limits.
What is lip-sync dubbing?
It is dubbing in which the new speech matches the speaker's mouth, and there are two ways to get there. Traditional lip-sync dubbing fits the words to the picture: the translation is written and performed to match the lip movements already on screen, and the picture is left alone. AI lip-sync dubbing fits the picture to the words: the new speech comes first, and the mouth is generated to match it, so the translation doesn't have to be bent around the old lip movements.
Which one does my project need?
Start from what is on screen. If no one's mouth is visible, a dub is enough. If a speaking face fills the frame, the words and the mouth have to agree, which takes lip-sync dubbing.
| Dubbing | Lip sync | AI lip-sync dubbing | |
|---|---|---|---|
| What changes | The speech | The mouth movement | Both |
| A speaking face on screen | Keeps moving to the old words | Moves with the audio you give it | Moves with the new words |
| Fits | Narration, screen recordings, faces too small to read | A still photo or an avatar that needs to speak | A speaker's face that fills the frame |
| With Sume | A transcript, your translation, TTS 1.0, then Timeline 1.0 | VEED Fabric 1.0 | Your translation, TTS 1.0, then VEED Fabric 1.0 |
How do I do each with Sume?
Each has a step-by-step post. To replace the speech under a video already in your Sume workspace, Translate a video's voiceover by API goes from transcript to translation to TTS 1.0 to a Timeline 1.0 render. In current code that render plays only its audio spine and an optional soundtrack, so the original voice is dropped, along with any music or room sound the video had.
To lip-sync dub an avatar video, How to dub an avatar video speaks the new lines with TTS 1.0 and animates the same avatar to them with Fabric. The Avatar 1.0 talking video can't re-voice itself: in current code it speaks English only.
Sources
Related posts
More in Sume Avatar 1.0
- How long should a microlearning video be? Length and AI
No fixed standard: one objective per video, and research on course videos favors 6 minutes or less. How to size lessons and make them with an AI avatar.
- Real-time AI avatar vs video avatar API: the difference
A real-time AI avatar talks live in a session; a video avatar API renders a finished file from a script. Which vendors sell which, and where Sume fits.
- Runway Characters API: avatars, live sessions and videos
Runway Characters builds avatars from one image, a voice and a personality, for live sessions or rendered videos. Endpoints, limits and pricing.
- Synthesia alternatives with an API: access, pricing, limits
Argil, Creatify, HeyGen and Sume each document an avatar video API. They differ in how API access is sold, the price unit and the length limits.
Written by Sume