Kandinsky 6.0 image-to-audio-video vs a first frame on Sume
Kandinsky 6.0 animates a still with synced audio. On Sume, frame_images starts a clip from your image, and some ids also take an audio reference. What fits.

To animate one still with sound, Kandinsky 6.0 Video has an image-to-audio-video mode; on Sume you send the still as a frame_images entry with frame_type: first_frame to a model that lists that type, and sound comes from the model or from an audio reference on ids that accept one. Sume does not list Kandinsky, so these are two routes to a similar result, not the same model.
Kandinsky facts are from its repository and technical report, read 2026-10-11. Sume facts are from Video Generation and Video Router.
What each side calls it
Kandinsky's repository names two modes, text-to-audio-video and image-to-audio-video, and the report uses the same pair. Both produce a 5-second clip with 44 kHz audio. Sume splits the image idea into two fields. frame_images gives first or last frame images and starts image-to-video. input_references gives style or content guidance and starts reference-to-video. If you send both, frame_images wins and the request runs as image-to-video.
| Sume id | Image start | Audio | Audio reference input |
|---|---|---|---|
| seedance-2 | first_frame and last_frame | Model can generate (generate_audio true in the docs example) | Accepts audio, video and image references |
| wan-3.0 | Check supported_frame_images | Check generate_audio | Accepts audio and video references |
| minimax-h3-max | First/last-frame image-to-video | Native stereo audio | Accepts audio and video references |
| gemini-omni-flash-1.1 | image_url plus optional end_image_url on Video Router | Always on | No: image and video references only |
Where the sound comes from
On Kandinsky the audio is generated from your description alongside the picture, so you describe dialogue, music or ambience in the prompt. On Sume, generate_audio tells the model to make audio or not, and the default is the model's own capability. Gemini Omni Flash 1.1 rejects generate_audio: false because its audio is always on.
Ids that accept audio_url references, namely the Seedance 2.x family, Wan 3.0, MiniMax H3 and H3 Max, can take a sound or a voice track as guidance. That is a closer fit than a prompt alone if you already have a sound design. What the docs do not say is how strictly any model follows the reference, so judge the output by ear.
Get the still ready
Image quality decides most of what an animated still looks like, so check three things before you submit.
- Send the image over public HTTPS in a supported format; the troubleshooting notes in the docs name unreachable or unsupported images as a cause of failed jobs.
- Match the aspect ratio of the still to the request. The video catalog lists the ratios each id accepts.
- Keep faces large and unobstructed if you want speech to look like it comes from them. Sume's own guidance is that video models do not lip-sync to generated TTS or to a later voice-over; for a face reading a script you already recorded, use the lip-sync surfaces instead.
Choose between them
Pick Kandinsky when you want an open pipeline you can alter and a 5-second talking beat where the picture and sound are generated together. Pick a Sume id when you want a job API, a clip longer than 5 seconds, or 1080p and up without a GPU. Run both on one still if the budget allows and compare, since neither vendor page offers a head-to-head.
Sources
Related posts
More in Comparisons
- Kandinsky 6.0 Video or a hosted API: a 5-second clip with sound
Kandinsky 6.0 Video is MIT open weights for 5 s clips with synced audio. Sume does not list it; two Sume ids make sound natively. Prices and limits.
- Kandinsky 6.0 stops at 5 s: a 30 s ad as one job or six clips
Kandinsky 6.0 Video makes 5 s clips. A 30 s ad on Sume is one Wan 3.0 job at $3.75 or six clips stitched for $0.17 in assembly fees. What each route costs you.
- Kandinsky 6.0 Lite takes 406 s on an RTX 5090: local or hosted?
The Kandinsky repo times one 5 s Full HD clip per GPU: 406 s for Lite on an RTX 5090. Clips per hour by card, and when a hosted Sume job is simpler.
- OpenRouter lists 26 video ids: which have a Sume id (Oct 2026)
Of 26 video ids on OpenRouter's listing, 11 map to a Sume catalog id and 15 do not. The full id-by-id table, read 2026-10-11, and what Sume lists instead.
Written by Sume