Kandinsky 6.0 lip-syncs in one pass; on Sume a talking face is Fabric

Kandinsky 6.0 Video makes speech and lip-sync inside the clip. Sume's docs send on-camera speech to Fabric or H3 Max Lip Sync with your audio, not a video id.

5 min readSume
All posts

Kandinsky 6.0 Video generates speech and lip-sync together with the picture in one pass; Sume does not use a video model for that job. Sume's models overview says video models do not lip-sync to generated TTS or to a later voice-over, and that every on-camera speaking shot goes through VEED Fabric 1.0 with an accepted still and audio.

So the difference is architectural: on the vendor pages Kandinsky makes the sound and the mouth movement in the same model, while Sume's route starts from a voice you already have. Which one you want depends on whether the script is fixed.

What Kandinsky does

The Kandinsky Lab page and the repository, both read on 2026-10-11, describe 5-second clips at 24 fps with synchronized 44 kHz audio and lip-sync. The modes are text-to-audio-video and image-to-audio-video. The vendor pages describe the sound as generated with the clip, and the inputs they list are a prompt and, for image mode, an image, so no audio file is part of the request.

What Sume's docs prescribe

The overview page lays out the speaking-shot rule and two audio-driven routes. Both take your audio and a still, so the script and voice are fixed before the video job starts.

Talking-face routes on Sume, from the models overview (read 2026-10-11)
RouteEndpointWhat you sendDocumented limit
VEED Fabric 1.0POST /v1/veed/fabric-1.0audio_url, measured duration_seconds, one visual sourceone image_url or avatar_handle, not both
MiniMax H3 Max Lip SyncPOST /v1/minimax/h3-max/lip-syncsame still and audio body as Fabricaudio 5 to 14.8 s, list x 1.25
Video ids (seedance, wan, kling...)POST /v1/videosprompt, optional framesnot for speaking shots per the docs

Fixed script or open script

If a legal line, a product name or a translated voice-over has to be exact, the Fabric route fits, because the audio file defines every syllable and the model animates a mouth to it. If you want a short, loose clip where a character says something plausible and the exact words do not matter, an audio-video model like Kandinsky is built for that brief, though you do not control the exact words.

Sume's catalog does list video ids with native sound, such as gemini-omni-flash-1.1 (audio always on) and the Seedance family (audio supported). Use them for ambience, effects and wordless motion. The docs say a face that talks is never a video-model clip with narration under it, so this page does not recommend them for dialogue.

Practical checks before you pick

A few questions settle most cases.

  • Do you have the audio already? If yes, the Fabric or H3 Max route takes it as audio_url. If no, produce it first, then send it.
  • How long is the line? H3 Max Lip Sync is documented for audio of 5 to 14.8 seconds; Kandinsky's clip is a fixed 5 seconds.
  • Do you need your own license terms? Kandinsky is MIT. A hosted route is an API contract, not a license to the weights.
  • Is the face yours? Both routes animate a still or generate a face; make sure you hold the rights to the person or character in it.

Bottom line

For a spoken line with an exact script, build the audio first and use a Sume talking-face route. For an open-ended 5-second scene with sound baked in, an open audio-video model is the closer match, and Sume does not carry Kandinsky. For pricing of the Sume routes per second, the existing post on lip sync API cost for H3 Max and Fabric has the table.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume