Kandinsky 6.0 lip-syncs in one pass; on Sume a talking face is Fabric
Kandinsky 6.0 Video makes speech and lip-sync inside the clip. Sume's docs send on-camera speech to Fabric or H3 Max Lip Sync with your audio, not a video id.

Kandinsky 6.0 Video generates speech and lip-sync together with the picture in one pass; Sume does not use a video model for that job. Sume's models overview says video models do not lip-sync to generated TTS or to a later voice-over, and that every on-camera speaking shot goes through VEED Fabric 1.0 with an accepted still and audio.
So the difference is architectural: on the vendor pages Kandinsky makes the sound and the mouth movement in the same model, while Sume's route starts from a voice you already have. Which one you want depends on whether the script is fixed.
What Kandinsky does
The Kandinsky Lab page and the repository, both read on 2026-10-11, describe 5-second clips at 24 fps with synchronized 44 kHz audio and lip-sync. The modes are text-to-audio-video and image-to-audio-video. The vendor pages describe the sound as generated with the clip, and the inputs they list are a prompt and, for image mode, an image, so no audio file is part of the request.
What Sume's docs prescribe
The overview page lays out the speaking-shot rule and two audio-driven routes. Both take your audio and a still, so the script and voice are fixed before the video job starts.
| Route | Endpoint | What you send | Documented limit |
|---|---|---|---|
| VEED Fabric 1.0 | POST /v1/veed/fabric-1.0 | audio_url, measured duration_seconds, one visual source | one image_url or avatar_handle, not both |
| MiniMax H3 Max Lip Sync | POST /v1/minimax/h3-max/lip-sync | same still and audio body as Fabric | audio 5 to 14.8 s, list x 1.25 |
| Video ids (seedance, wan, kling...) | POST /v1/videos | prompt, optional frames | not for speaking shots per the docs |
Fixed script or open script
If a legal line, a product name or a translated voice-over has to be exact, the Fabric route fits, because the audio file defines every syllable and the model animates a mouth to it. If you want a short, loose clip where a character says something plausible and the exact words do not matter, an audio-video model like Kandinsky is built for that brief, though you do not control the exact words.
Sume's catalog does list video ids with native sound, such as gemini-omni-flash-1.1 (audio always on) and the Seedance family (audio supported). Use them for ambience, effects and wordless motion. The docs say a face that talks is never a video-model clip with narration under it, so this page does not recommend them for dialogue.
Practical checks before you pick
A few questions settle most cases.
- Do you have the audio already? If yes, the Fabric or H3 Max route takes it as
audio_url. If no, produce it first, then send it. - How long is the line? H3 Max Lip Sync is documented for audio of 5 to 14.8 seconds; Kandinsky's clip is a fixed 5 seconds.
- Do you need your own license terms? Kandinsky is MIT. A hosted route is an API contract, not a license to the weights.
- Is the face yours? Both routes animate a still or generate a face; make sure you hold the rights to the person or character in it.
Bottom line
For a spoken line with an exact script, build the audio first and use a Sume talking-face route. For an open-ended 5-second scene with sound baked in, an open audio-video model is the closer match, and Sume does not carry Kandinsky. For pricing of the Sume routes per second, the existing post on lip sync API cost for H3 Max and Fabric has the table.
Sources
Related posts
More in Comparisons
- Kandinsky 6.0 Pro: 292 s a clip on an H100, 100 clips take 8.1 hours
Kandinsky's published timing is 292 seconds per 5-second Pro HD clip on an H100. That is 8.1 hours for 100 clips, set against hosted per-clip prices on Sume.
- Is Kandinsky 6.0 Video on Sume? No, and the closest audio ids
Kandinsky 6.0 Video is open weights with no hosted API, and Sume lists no Kandinsky id. These Sume ids give 5-second clips with sound and a first frame.
- Kandinsky 6.0 image-to-audio-video vs a first frame on Sume
Kandinsky 6.0 animates a still with synced audio. On Sume, frame_images starts a clip from your image, and some ids also take an audio reference. What fits.
- Kandinsky 6.0 Video or a hosted API: a 5-second clip with sound
Kandinsky 6.0 Video is MIT open weights for 5 s clips with synced audio. Sume does not list it; two Sume ids make sound natively. Prices and limits.
Written by Sume