Lip sync for singing: fal can turn off guidance, Sume has no switch
fal's H3 Max Lip Sync is transcription-guided by default and can be switched off for singing or processed audio. Sume's documented body has no such field.

On fal, H3 Max Lip Sync uses transcription-guided sync by default and the page says you can disable it for singing or processed audio. Sume's POST /v1/minimax/h3-max/lip-sync documents no switch for that: its body is a still or avatar, a Sume-hosted audio_url, duration_seconds, resolution, an ignored speed_tier, and mode. For a song, expect speech-style sync and test first.
What fal says
fal's page describes the model as turning a single image and an audio file into a talking video, post-trained by fal from MiniMax H3. Audio must be 5 to 15 seconds, output length matches the audio, and longer audio is clipped. It says the model supports any language and takes photographs, 3D renders, illustrations and paintings with a visible face, at an aspect ratio between 0.4 and 2.5.
The transcription-guided default is the interesting line for singers. Sung words are stretched and sit over music, so a model that listens for words can mislead. fal's answer is a toggle.
What Sume's body exposes
Sume's lip-sync request mirrors Fabric's: exactly one visual source (image_url or avatar_id or avatar_handle), audio_url, a duration_seconds between 5 and 14.8, a resolution of 480p, 768p (default) or 1080p, speed_tier (accepted and ignored) and mode. It also rejects model, endpoint and provider_endpoint, because the URL is the model.
I found no field in those docs for turning guidance off, so I do not claim Sume forwards one.
| Item | fal H3 Max Lip Sync | Sume lip-sync route |
|---|---|---|
| Transcription guidance | On by default, can be disabled | No documented field |
| Audio window | 5 to 15 s | 5 to 14.8 s, never clamped |
| Image types | Photo, 3D, illustration, painting | Public HTTPS still, aspect 0.4 to 2.5 |
| Resolutions | 480p to 2K | 480p, 768p, 1080p; no 2K |
| Price per second at 768p | $0.08 list | $0.08 x 1.25 = $0.10 |
How to test a song clip
Cut a 6 to 10 second chorus and render it once at 480p, which is $0.0625 a second, so a 10-second test costs $0.625. Watch the mouth against the beat. If the sync fails, change the audio instead: use an a cappella take with less music under it.
For talking audio, the route works as documented. For sung audio, treat it as an experiment and keep the clip short.
Sources
Related posts
More in Developers
- HEAD-check reference image URLs before an image edit call (Python)
OpenAI caps edit images at 50 MB and Ideogram at 25 MB. A Python pre-flight that checks size, type and https before you send references to Sume.
- Hold AI Shorts for human approval: poll the run, then publish
A Python poll loop for a Sume Format run that stops at a human yes or no before any upload, so a reviewer can reject sameness first.
- How long can your webhook be down? Sume job vs run retry windows
Sume job webhooks retry 10 times, 30 s apart: about 5 minutes. Run webhooks back off to roughly 3 hours. The arithmetic, and what Redeliver covers.
- How long does AI lip sync take? fal says about a minute at 1080p
fal says a 1080p H3 Max lip-sync clip takes about a minute. On Sume the call is an async job, so poll with backoff. This Python example submits and polls.
Written by Sume