Kandinsky 6.0 video with sound: is it on Sume, and what to use

Kandinsky 6.0 makes video and audio together, MIT-licensed. Sume does not list it. Here is what Sume lists for synced sound, and when to self-host Kandinsky.

4 min readSume
All posts

No, Sume does not list Kandinsky 6.0 video. The Sume API serves video with synchronized sound from other models, and you can read the live list from GET /v1/videos/models instead of trusting a blog post.

Kandinsky 6.0 is a family of open models that generate a video and its audio in one pass. It was released on 6 October 2026 under the MIT license, so the real question for most teams is not whether a hosted API exists, but whether to run it themselves or buy sound-capable clips from a hosted catalog.

What Kandinsky 6.0 is, from the vendor pages

The technical report describes two sizes, a 3B Lite model and a 29B Pro model. Both generate 5-second clips with synchronized 44 kHz audio, including lip-sync, and both support text-to-audio-video and image-to-audio-video modes. The base output is standard definition, and a built-in super-resolution stage lifts it to Full HD (1920x1080). Code and checkpoints are MIT-licensed, and the report describes consumer-GPU deployment through block offloading and VRAM presets.

The GitHub repository lists the checkpoints it ships: pro, pro-pretrain, lite, lite-distill and lite-pretrain, with a base resolution of 864x480. It names RTX 4090, RTX 5090, A100 80GB and H100 as example cards. Neither the repository nor the project site gives an API price, which is why there is nothing to compare against Sume's per-second rates.

Kandinsky 6.0 facts as published by the vendor pages (read 2026-10-11)
ItemWhat the vendor pages say
SizesLite 3B, Pro 29B
Clip length5 seconds
AudioSynchronized 44 kHz, including lip-sync
ModesText-to-audio-video and image-to-audio-video
Base output864x480, Full HD through a super-resolution stage
LicenseMIT for code and checkpoints
Hosted API priceNot published on the pages read

What Sume lists for video with built-in audio

Sume's video catalog is the source of truth, and it changes. As of this writing the catalog code marks these routes as producing audio with the picture. Check the live list before you build on any of them.

  • seedance-2.5: 4 to 30 seconds, 480p/720p/1080p, native audio, and it accepts reference audio.
  • wan-3.0: 2 to 30 seconds, native audio, 480p/720p/1080p.
  • minimax-h3: 5 to 15 seconds, always native stereo audio.
  • gemini-omni-flash-1.1: 3 to 10 seconds, native audio, but no audio reference.
  • kling-3: 4 to 15 seconds at 1080p with an audio toggle.
Sume video routes with native audio, from the Sume video catalog code and https://docs.sume.com/models/videos, read 2026-10-11
Model idLength rangeAudio behavior
seedance-2.54-30 sNative audio, reference audio accepted
wan-3.02-30 sNative audio
minimax-h35-15 sAlways native stereo audio
gemini-omni-flash-1.13-10 sNative audio, no audio reference
kling-34-15 sAudio toggle

Which path fits which job

Pick Sume when you want a hosted API, a job id you can poll, a webhook, and per-second billing with no GPU to babysit. Pick a self-hosted Kandinsky when you need open weights, offline generation or full control of the checkpoint, and you have a card in the class the repository names.

Length is the practical difference. Kandinsky 6.0 stops at 5 seconds per clip, according to the vendor pages. Sume's Seedance 2.5 and Wan 3.0 routes accept up to 30 seconds in one request, so a single spoken beat does not have to be stitched from several generations.

If your need is an exact spoken script rather than a generated voice, that is a separate job. Sume's docs say a speaking face is a lip-sync route (Fabric with an accepted still, or MiniMax H3 Max Lip Sync) fed with audio from the TTS endpoints, because video models do not lip-sync to TTS or voice-over you supply.

How to check the catalog yourself

Ask the API what it serves today and look for the id you need. This is a read-only call and costs nothing.

import os, json, urllib.request

req = urllib.request.Request(
    "https://api.sume.com/v1/videos/models",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
)
with urllib.request.urlopen(req) as resp:
    body = json.load(resp)

print(json.dumps(body, indent=2)[:2000])

What to do next

If you only need a sound-capable clip today, start from the table above and the video generation docs. If Kandinsky is the model you want, self-host it from the repository and treat Sume as the place to generate longer shots, add TTS narration, or render lip-sync from a script. The API reference covers authentication and job polling for the Sume side.

Sources

Related posts

More in Models

All Models posts

Written by Sume