Kandinsky 6.0 Video is MIT: can you use the clips commercially?
Kandinsky 6.0 Video Pro (29B) and Lite (3B) are MIT-licensed with joint audio. What MIT covers, what to check, and the hosted alternative on Sume.

Yes on the headline question: Kandinsky 6.0 Video is published with its code and model checkpoints under the MIT license, per its paper and repository, and MIT has no revenue cap, territory list or attribution banner of the kind some other open video licenses carry. Read the notice file and any third-party component licenses in the repo before you ship, because the MIT text covers the project, not every dependency it pulls in.
What exactly is released?
The paper is dated October 6, 2026. It describes two models that generate picture and sound together from text or from an image: Kandinsky 6.0 Video Lite at 3 billion parameters and Pro at 29 billion. Clips are 5 seconds long. The base video is generated at SD resolution and upscaled to Full HD (1920 by 1080) by a separate super-resolution model, so "Full HD" is an upscale step, not native output.
The repository lists integrations with Diffusers, ComfyUI, vLLM-Omni, SGLang and FastVideo.
| Item | Lite | Pro |
|---|---|---|
| Parameters | 3B | 29B |
| Clip length | 5 seconds | 5 seconds |
| Audio | Synchronized 44 kHz, with lip-sync | Synchronized 44 kHz, with lip-sync |
| Modes | Text and image to audio-video | Text and image to audio-video |
| License | MIT | MIT |
What does MIT give you, and what does it not?
MIT permits use, copying, modification and distribution of the software, including commercially, provided the copyright and permission notice travel with copies. It carries no warranty. It does not say anything about whether a particular clip infringes someone's likeness or trademark, and it does not replace platform rules or disclosure laws for AI-generated media in your market.
I did not audit the licenses of the text encoders or other components the install pulls in. Check them the way you would for any open-weights stack.
What does it cost to run yourself?
The repository gives one concrete figure: a Pro HD video takes about 292 seconds on an H100 and about 2,716 seconds on an RTX 5080. Setup needs Python 3.13 or 3.14 and installs a PyTorch build matched to your GPU. The paper mentions consumer-GPU deployment with block offloading and VRAM presets but the pages I read give no exact VRAM floor, so measure on your card before planning a batch.
Is it on Sume, and what is the practical choice?
I searched the Sume repository and found no Kandinsky model, so do not send it as a model id. Sume's Video Router docs list other rows: wan-3.0 accepts 2 to 30 seconds, minimax-h3 accepts 5 to 15 seconds at 480p or 768p, and gemini-omni-flash-1.1 is 3 to 10 seconds with native synced audio always on.
If you need exactly Kandinsky's output or want to fine-tune, run the MIT weights yourself. If you need a 7-second or 30-second clip, or no GPU to babysit, pick a hosted row and read its capabilities from the catalog first.
Sources
Related posts
More in Models
- Kimi K3 on Sume: no catalog row, and what to pick for a video agent
Kimi K3 is not in Sume's agent model catalog. Moonshot says it takes text, images and video; here is the verified alternative and how Kimi can still call Sume.
- Korean clip? Whistle's seven languages vs Sume STT
Cactus Whistle lists English, German, French, Spanish, Italian, Dutch and Polish. For a Korean clip use Sume STT with a ko hint, then caption it.
- Muse Spark 1.3 on Sume: picker row, tool use and media jobs
Meta says Muse Spark 1.3 uses about 20% fewer tool calls. Sume lists it as a catalog row behind the OpenRouter switch; the API cannot pick it.
- Nova 2.5 Sonic for a narration file? Speech-to-speech vs Sume TTS
Nova 2.5 Sonic is built for live voice agents. For a finished narration file, Sume TTS takes a transcript up to 20,000 characters and returns audio as a job.
Written by Sume