Lip sync with a hand over the mouth: sync-3 vs Fabric on Sume

Sync Labs says sync-3 handles obstructions on faces. Sume's lip-sync routes start from a still plus audio. What each takes as input and what is promised.

4 min readSume
All posts

Sync Labs describes sync-3 as handling close-ups, extreme angles and obstructions automatically. Sume's lip-sync comes from Fabric, which takes a still image plus audio and returns a talking clip. The Sume docs make no promise about a hand over the mouth, so pick a clear still instead of relying on the model to cope.

What the summary says

These points come from the Sync Labs docs introduction. The page does not say how well sync-3 copes with a hand over the mouth, so test it on your own footage.

Sync Labs docs (read 2026-10-03)
ItemReported
sync-3Native 4K; handles close-ups, extreme angles and obstructions automatically
PriceListed from $0.02 per second (lipsync-1.9) to $0.133 per second (sync-3)

What Sume's routes take

Sume lists two still-plus-audio lip-sync routes. VEED Fabric 1.0 is POST /v1/veed/fabric-1.0, with the model id veed/fabric-1.0. MiniMax H3 Max Lip Sync is POST /v1/minimax/h3-max/lip-sync with the same still-and-audio body, audio from 5 to 14.8 seconds, billed at list times 1.25.

  • Send audio_url, a measured duration_seconds, and exactly one visual source.
  • The visual source is image_url or avatar_handle, not both.
  • Prefer a generated, inspected, posed still. Use avatar_handle only when the user named that avatar.

What this means for occlusion

Because the input is a still you choose, the way to avoid a blocked mouth is to inspect the still and reject one where a hand, microphone or product covers the face. Sume's docs do not describe how Fabric behaves when the mouth is covered, so treat it as unsupported and test.

Also note the shape difference. These Sume routes build a talking clip from a still, not from an existing video. If you need to re-voice footage that is already shot, they are a different job from what a video-to-video lip-sync tool does.

Where video models fit

Sume's models guide states that video models do not lip-sync to generated TTS or to a later voice-over. A talking face is therefore a Fabric clip with an accepted still and the speech audio, and wordless beats, B-roll and product motion go to a video model. Check price per second in the live catalog before you compare it with the $0.133 per second Sync Labs lists for sync-3.

Sources

Related posts

More in Models

All Models posts

Written by Sume