Lip sync with a hand over the mouth: sync-3 vs Fabric on Sume
Sync Labs says sync-3 handles obstructions on faces. Sume's lip-sync routes start from a still plus audio. What each takes as input and what is promised.

Sync Labs describes sync-3 as handling close-ups, extreme angles and obstructions automatically. Sume's lip-sync comes from Fabric, which takes a still image plus audio and returns a talking clip. The Sume docs make no promise about a hand over the mouth, so pick a clear still instead of relying on the model to cope.
What the summary says
These points come from the Sync Labs docs introduction. The page does not say how well sync-3 copes with a hand over the mouth, so test it on your own footage.
| Item | Reported |
|---|---|
| sync-3 | Native 4K; handles close-ups, extreme angles and obstructions automatically |
| Price | Listed from $0.02 per second (lipsync-1.9) to $0.133 per second (sync-3) |
What Sume's routes take
Sume lists two still-plus-audio lip-sync routes. VEED Fabric 1.0 is POST /v1/veed/fabric-1.0, with the model id veed/fabric-1.0. MiniMax H3 Max Lip Sync is POST /v1/minimax/h3-max/lip-sync with the same still-and-audio body, audio from 5 to 14.8 seconds, billed at list times 1.25.
- Send
audio_url, a measuredduration_seconds, and exactly one visual source. - The visual source is
image_urloravatar_handle, not both. - Prefer a generated, inspected, posed still. Use
avatar_handleonly when the user named that avatar.
What this means for occlusion
Because the input is a still you choose, the way to avoid a blocked mouth is to inspect the still and reject one where a hand, microphone or product covers the face. Sume's docs do not describe how Fabric behaves when the mouth is covered, so treat it as unsupported and test.
Also note the shape difference. These Sume routes build a talking clip from a still, not from an existing video. If you need to re-voice footage that is already shot, they are a different job from what a video-to-video lip-sync tool does.
Where video models fit
Sume's models guide states that video models do not lip-sync to generated TTS or to a later voice-over. A talking face is therefore a Fabric clip with an accepted still and the speech audio, and wordless beats, B-roll and product motion go to a video model. Check price per second in the live catalog before you compare it with the $0.133 per second Sync Labs lists for sync-3.
Sources
Related posts
More in Models
- Luma's 2026 timeline: Ray3.14, Ray3.2, Scenes and Variants
Luma shipped Ray3.14 in January, Ray3.2 in June, Scenes in August and Variants on Oct 1, 2026. What each added, and why to pin model ids.
- Lyria 3.5 blocks artist-voice prompts: how to write briefs that pass
Google's Lyria 3.5 docs note that prompts asking for specific artist voices are blocked. Describe the sound instead, then run it through the Sume Music Router.
- MAI-Voice-2.1 emotion control vs Sume's emotion field
Microsoft lists emotion control on MAI-Voice-2.1. Sume's TTS takes a free-text emotion string, speed and volume. What each gives you, and how to test it.
- MiniMax H3 Max: the prompt-adherence variant on Sume
fal describes MiniMax H3 Max as tuned for prompt adherence. On Sume, minimax-h3-max runs 480p to 1080p for 5 to 15 s with frames and references.
Written by Sume