Will an AI video model lip-sync a voice-over added afterwards?

No: Sume's docs say video models do not lip-sync to generated speech or a later voice-over. A talking face needs a still plus audio, or an avatar video.

5 min readSume
All posts

No. Sume's models documentation states that video models do not lip-sync to generated text-to-speech or to a voice-over added later, so a face that talks is never a video-model clip with narration under it. For a person speaking on camera, use a talking-still route or an avatar video, and keep video models for wordless beats.

The rule, as the docs state it

The Models overview (read 2026-10-08) sets out the split. Every shot where a person speaks on camera is a Fabric shot: an accepted still plus TTS audio. That applies to short UGC and presenter ads, testimonials, recreated beats where someone speaks, and live-commerce host talk. Wordless beats, B-roll and product motion go through image then video, with no Fabric step.

Which route fits each shot, as of 2026-10-08
Shot typeRouteSource
Person speaks to cameraTalking still plus audio (veed/fabric-1.0) or Avatar 1.0 talking-videoModels overview
Product motion, B-rollImage, inspect, then videoModels overview
Wordless character beatVideo model without narration claimsModels overview

What goes wrong if you ignore it

If you generate a silent clip of a person and lay narration over it, the mouth will not follow the words. At best the speaker looks like they are doing a dub; at worst the mismatch is obvious in the first second. Editing cannot fix a mismatch that is built into the footage.

The docs' fix is to decide, shot by shot, whether anyone speaks. Mark those shots as speech shots up front and generate them through a route that takes audio or a script as input.

Two ways to get a talking face on Sume

  • Talking still: POST /v1/veed/fabric-1.0 with an audio_url, a measured duration_seconds and one visual source (an image_url of an inspected still, or an avatar_handle the user named). You cannot send both visual sources.
  • Avatar video: POST /v1/avatar-1.0/talking-video with an avatar_handle and a script or video_inputs, 4-60 seconds, so Sume voices the script itself.
  • A MiniMax H3 Max lip-sync route takes the same still plus audio body, with audio of 5-14.8 seconds, per the models page.

Practical checklist

Write the script first and split lines into speech and no-speech beats. Generate a still of the speaker and inspect it. Send speech beats to the talking route and everything else to image or video. Join the pieces afterwards on the timeline. Media URLs must be public HTTPS; see Media inputs.

What to check in your own pipeline

Audit every shot list for lines of dialogue. For each, ask whether a face is visible while the words are spoken. If the answer is yes, that shot needs a speech route. If the face is off screen, or the shot is a product close-up, a video model with narration over it is fine, because there is no mouth to mismatch.

Also keep timing in mind. A talking-still job needs a measured duration_seconds that matches the audio. Measure the file rather than estimating, and keep the audio and the requested duration aligned so the clip does not run short or leave dead frames at the end.

When you assemble the final cut, treat speech shots as fixed. Cut around them with B-roll instead of re-timing the spoken clip, because trimming a lip-synced take mid-word breaks the match. Timeline and trim utilities exist in the Sume docs for the assembly step.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume