Will an AI video model lip-sync a voice-over added afterwards?
No: Sume's docs say video models do not lip-sync to generated speech or a later voice-over. A talking face needs a still plus audio, or an avatar video.

No. Sume's models documentation states that video models do not lip-sync to generated text-to-speech or to a voice-over added later, so a face that talks is never a video-model clip with narration under it. For a person speaking on camera, use a talking-still route or an avatar video, and keep video models for wordless beats.
The rule, as the docs state it
The Models overview (read 2026-10-08) sets out the split. Every shot where a person speaks on camera is a Fabric shot: an accepted still plus TTS audio. That applies to short UGC and presenter ads, testimonials, recreated beats where someone speaks, and live-commerce host talk. Wordless beats, B-roll and product motion go through image then video, with no Fabric step.
| Shot type | Route | Source |
|---|---|---|
| Person speaks to camera | Talking still plus audio (veed/fabric-1.0) or Avatar 1.0 talking-video | Models overview |
| Product motion, B-roll | Image, inspect, then video | Models overview |
| Wordless character beat | Video model without narration claims | Models overview |
What goes wrong if you ignore it
If you generate a silent clip of a person and lay narration over it, the mouth will not follow the words. At best the speaker looks like they are doing a dub; at worst the mismatch is obvious in the first second. Editing cannot fix a mismatch that is built into the footage.
The docs' fix is to decide, shot by shot, whether anyone speaks. Mark those shots as speech shots up front and generate them through a route that takes audio or a script as input.
Two ways to get a talking face on Sume
- Talking still: POST /v1/veed/fabric-1.0 with an audio_url, a measured duration_seconds and one visual source (an image_url of an inspected still, or an avatar_handle the user named). You cannot send both visual sources.
- Avatar video: POST /v1/avatar-1.0/talking-video with an avatar_handle and a script or video_inputs, 4-60 seconds, so Sume voices the script itself.
- A MiniMax H3 Max lip-sync route takes the same still plus audio body, with audio of 5-14.8 seconds, per the models page.
Practical checklist
Write the script first and split lines into speech and no-speech beats. Generate a still of the speaker and inspect it. Send speech beats to the talking route and everything else to image or video. Join the pieces afterwards on the timeline. Media URLs must be public HTTPS; see Media inputs.
What to check in your own pipeline
Audit every shot list for lines of dialogue. For each, ask whether a face is visible while the words are spoken. If the answer is yes, that shot needs a speech route. If the face is off screen, or the shot is a product close-up, a video model with narration over it is fine, because there is no mouth to mismatch.
Also keep timing in mind. A talking-still job needs a measured duration_seconds that matches the audio. Measure the file rather than estimating, and keep the audio and the requested duration aligned so the clip does not run short or leave dead frames at the end.
When you assemble the final cut, treat speech shots as fixed. Cut around them with B-roll instead of re-timing the spoken clip, because trimming a lip-synced take mid-word breaks the match. Timeline and trim utilities exist in the Sume docs for the assembly step.
Sources
Related posts
More in Sume Avatar 1.0
- Face-swap clip on YouTube: label needed? Sume's beta caps at 15 s
YouTube asks for a label when a real person appears to say or do something they did not. Sume's face swap beta takes about 4 to 15 seconds; $3.68 max at plus.
- Face swap or avatar video? Pick by whether you already have footage
Sume's face swap Beta applies an avatar face to your own 4-15 second video; Avatar Video renders a new clip from a script, up to 60 seconds. How to choose.
- Five 20-word avatar sentences make five 8-second clips, not three
Why five 20-word sentences plan as five 8-second clips (40 s) in Avatar 1.0, and how pairing shorter sentences saves clips.
- One wrong sentence in an avatar script: what 5 re-renders cost
Changing a script means a new Avatar 1.0 job. Five full re-renders of a 20-second take cost $18.40 on standard, $24.50 on plus and $55.00 on max.
Written by Sume