Why a voiceover on an AI video clip never lip-syncs

Video models do not lip-sync to TTS or a later voiceover. Per Sume's docs a talking face comes from Fabric, H3 Max Lip Sync or an avatar talking-video job.

5 min readSume
All posts

A video model clip with narration laid over it will not lip-sync, because video models do not lip-sync to generated TTS or to a later voiceover. Sume's models page states the rule: a face that talks is never a video-model clip with narration under it. Use a talking-face route, where the mouth is driven by the audio.

The three talking-face routes

Sume lists these on its models pages:

  • VEED Fabric 1.0 at POST /v1/veed/fabric-1.0: a still plus an audio_url.
  • MiniMax H3 Max Lip Sync at POST /v1/minimax/h3-max/lip-sync: the same still plus audio body, with audio of 5 to 14.8 seconds.
  • Avatar Video at POST /v1/avatar-1.0/talking-video: a ready avatar and a script, 4 to 60 seconds.

What belongs on a video model instead

Wordless beats, B-roll and product motion are for the video models. Generate those with a normal video model, and join them around the talking-face clip on a Timeline so the audio of the face clip stays in sync with its own picture.

Fabric's input rules

Send audio_url, a measured duration_seconds and exactly one visual source, either an image_url or an avatar. You cannot send both. Measure the audio first, because duration_seconds is the basis of the reservation.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume