VEED lipsync-v2 video + audio vs Sume still + audio lip sync
fal lists veed/lipsync-v2 as video plus audio in. Sume lip sync routes start from a still and an audio clip, so existing footage is not re-synced.

fal lists veed/lipsync-v2 as taking an existing video file plus an audio file. The Sume models overview lists no video-in lip-sync route; its lip-sync routes, VEED Fabric 1.0 and MiniMax H3 Max, start from a still image and an audio clip, so they create a new talking clip rather than re-syncing footage you already shot.
The vendor input types are from fal's veed/lipsync-v2 page, read 2026-10-01; the snapshot shows no usable price, so none is quoted. Sume facts are from the Models overview.
What inputs does veed/lipsync-v2 take on fal?
The page's upload hints list video as mp4, mov, webm, m4v or gif, and audio as mp3, ogg, wav, m4a or aac.
| Route | Visual input | Audio input |
|---|---|---|
fal veed/lipsync-v2 | Video: mp4, mov, webm, m4v, gif | mp3, ogg, wav, m4a, aac |
Sume veed/fabric-1.0 | Still image (image_url or avatar_handle) | audio_url |
Sume minimax/h3-max/lip-sync | Still image | Audio 5 to 14.8 s |
Can Sume re-lip-sync a video I already have?
Not through the routes in the models overview; both take a still plus audio. The docs also say video models do not lip-sync to generated TTS or to a later voice-over, so laying new audio under a generated clip does not produce matched lips.
What is the closest workflow on Sume?
Pull a still of the speaker from the clip with POST /v1/video-frames, which is unbilled and takes up to 24 stills at chosen seconds, then send that still and your audio to Fabric. The result is a new talking clip of that face, not a repaired version of the original take. The clip must be on media.sume.com and no longer than 300 seconds for video frames.
When is each shape the right one?
If you have footage with the wrong or dubbed audio and need the original shot kept, a video-in route is what you want, and the models overview lists no such route. If you are building a talking presenter from a photo, the still-in routes fit; see lip-sync API with photo and audio.
Sources
Related posts
More in Models
- Veo 3.1 only makes 16:9 and 9:16; which Sume video models add more
Google's Veo 3.1 supports 16:9 and 9:16. On Sume, Seedance, MiniMax and Wan add 4:3, 1:1 and 3:4, Kling adds 1:1, and Grok takes no ratio.
- Veo 3.1 prompts cap at 1,024 tokens; Sume's Omni at 20,000 characters
Google caps a Veo 3.1 text prompt at 1,024 tokens. On Sume, gemini-omni-flash-1.1 documents a 20,000-character cap. Tokens and characters differ.
- Vidu Q2 Pro Fast image-to-video 1080p vs Sume first frame
QwenCloud lists vidu/viduq2-pro-fast_img2video at 720P and 1080P. Sume has no Vidu id; send image_url as the first frame to a 1080p model.
- Vidu Q3 ad reference-to-video vs Sume reference images
QwenCloud lists a Vidu Q3 ad reference-to-video model. Vidu is not in Sume's video ids; reference images work on models that list them.
Written by Sume