Kandinsky 6.0 image-to-audio-video vs a first frame on Sume

Kandinsky 6.0 animates a still with synced audio. On Sume, frame_images starts a clip from your image, and some ids also take an audio reference. What fits.

5 min readSume
All posts

To animate one still with sound, Kandinsky 6.0 Video has an image-to-audio-video mode; on Sume you send the still as a frame_images entry with frame_type: first_frame to a model that lists that type, and sound comes from the model or from an audio reference on ids that accept one. Sume does not list Kandinsky, so these are two routes to a similar result, not the same model.

Kandinsky facts are from its repository and technical report, read 2026-10-11. Sume facts are from Video Generation and Video Router.

What each side calls it

Kandinsky's repository names two modes, text-to-audio-video and image-to-audio-video, and the report uses the same pair. Both produce a 5-second clip with 44 kHz audio. Sume splits the image idea into two fields. frame_images gives first or last frame images and starts image-to-video. input_references gives style or content guidance and starts reference-to-video. If you send both, frame_images wins and the request runs as image-to-video.

Image inputs on Sume, from the Video Generation and Video Router docs (read 2026-10-11).
Sume idImage startAudioAudio reference input
seedance-2first_frame and last_frameModel can generate (generate_audio true in the docs example)Accepts audio, video and image references
wan-3.0Check supported_frame_imagesCheck generate_audioAccepts audio and video references
minimax-h3-maxFirst/last-frame image-to-videoNative stereo audioAccepts audio and video references
gemini-omni-flash-1.1image_url plus optional end_image_url on Video RouterAlways onNo: image and video references only

Where the sound comes from

On Kandinsky the audio is generated from your description alongside the picture, so you describe dialogue, music or ambience in the prompt. On Sume, generate_audio tells the model to make audio or not, and the default is the model's own capability. Gemini Omni Flash 1.1 rejects generate_audio: false because its audio is always on.

Ids that accept audio_url references, namely the Seedance 2.x family, Wan 3.0, MiniMax H3 and H3 Max, can take a sound or a voice track as guidance. That is a closer fit than a prompt alone if you already have a sound design. What the docs do not say is how strictly any model follows the reference, so judge the output by ear.

Get the still ready

Image quality decides most of what an animated still looks like, so check three things before you submit.

  • Send the image over public HTTPS in a supported format; the troubleshooting notes in the docs name unreachable or unsupported images as a cause of failed jobs.
  • Match the aspect ratio of the still to the request. The video catalog lists the ratios each id accepts.
  • Keep faces large and unobstructed if you want speech to look like it comes from them. Sume's own guidance is that video models do not lip-sync to generated TTS or to a later voice-over; for a face reading a script you already recorded, use the lip-sync surfaces instead.

Choose between them

Pick Kandinsky when you want an open pipeline you can alter and a 5-second talking beat where the picture and sound are generated together. Pick a Sume id when you want a job API, a clip longer than 5 seconds, or 1080p and up without a GPU. Run both on one still if the budget allows and compare, since neither vendor page offers a head-to-head.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume