Lip sync for singing: fal can turn off guidance, Sume has no switch

fal's H3 Max Lip Sync is transcription-guided by default and can be switched off for singing or processed audio. Sume's documented body has no such field.

5 min readSume
All posts

On fal, H3 Max Lip Sync uses transcription-guided sync by default and the page says you can disable it for singing or processed audio. Sume's POST /v1/minimax/h3-max/lip-sync documents no switch for that: its body is a still or avatar, a Sume-hosted audio_url, duration_seconds, resolution, an ignored speed_tier, and mode. For a song, expect speech-style sync and test first.

What fal says

fal's page describes the model as turning a single image and an audio file into a talking video, post-trained by fal from MiniMax H3. Audio must be 5 to 15 seconds, output length matches the audio, and longer audio is clipped. It says the model supports any language and takes photographs, 3D renders, illustrations and paintings with a visible face, at an aspect ratio between 0.4 and 2.5.

The transcription-guided default is the interesting line for singers. Sung words are stretched and sit over music, so a model that listens for words can mislead. fal's answer is a toggle.

What Sume's body exposes

Sume's lip-sync request mirrors Fabric's: exactly one visual source (image_url or avatar_id or avatar_handle), audio_url, a duration_seconds between 5 and 14.8, a resolution of 480p, 768p (default) or 1080p, speed_tier (accepted and ignored) and mode. It also rejects model, endpoint and provider_endpoint, because the URL is the model.

I found no field in those docs for turning guidance off, so I do not claim Sume forwards one.

Singing and processed audio on two surfaces (read 2026-10-05)
Itemfal H3 Max Lip SyncSume lip-sync route
Transcription guidanceOn by default, can be disabledNo documented field
Audio window5 to 15 s5 to 14.8 s, never clamped
Image typesPhoto, 3D, illustration, paintingPublic HTTPS still, aspect 0.4 to 2.5
Resolutions480p to 2K480p, 768p, 1080p; no 2K
Price per second at 768p$0.08 list$0.08 x 1.25 = $0.10

How to test a song clip

Cut a 6 to 10 second chorus and render it once at 480p, which is $0.0625 a second, so a 10-second test costs $0.625. Watch the mouth against the beat. If the sync fails, change the audio instead: use an a cappella take with less music under it.

For talking audio, the route works as documented. For sung audio, treat it as an experiment and keep the clip short.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume