Kandinsky 6.0 video with sound: is it on Sume, and what to use
Kandinsky 6.0 makes video and audio together, MIT-licensed. Sume does not list it. Here is what Sume lists for synced sound, and when to self-host Kandinsky.

No, Sume does not list Kandinsky 6.0 video. The Sume API serves video with synchronized sound from other models, and you can read the live list from GET /v1/videos/models instead of trusting a blog post.
Kandinsky 6.0 is a family of open models that generate a video and its audio in one pass. It was released on 6 October 2026 under the MIT license, so the real question for most teams is not whether a hosted API exists, but whether to run it themselves or buy sound-capable clips from a hosted catalog.
What Kandinsky 6.0 is, from the vendor pages
The technical report describes two sizes, a 3B Lite model and a 29B Pro model. Both generate 5-second clips with synchronized 44 kHz audio, including lip-sync, and both support text-to-audio-video and image-to-audio-video modes. The base output is standard definition, and a built-in super-resolution stage lifts it to Full HD (1920x1080). Code and checkpoints are MIT-licensed, and the report describes consumer-GPU deployment through block offloading and VRAM presets.
The GitHub repository lists the checkpoints it ships: pro, pro-pretrain, lite, lite-distill and lite-pretrain, with a base resolution of 864x480. It names RTX 4090, RTX 5090, A100 80GB and H100 as example cards. Neither the repository nor the project site gives an API price, which is why there is nothing to compare against Sume's per-second rates.
| Item | What the vendor pages say |
|---|---|
| Sizes | Lite 3B, Pro 29B |
| Clip length | 5 seconds |
| Audio | Synchronized 44 kHz, including lip-sync |
| Modes | Text-to-audio-video and image-to-audio-video |
| Base output | 864x480, Full HD through a super-resolution stage |
| License | MIT for code and checkpoints |
| Hosted API price | Not published on the pages read |
What Sume lists for video with built-in audio
Sume's video catalog is the source of truth, and it changes. As of this writing the catalog code marks these routes as producing audio with the picture. Check the live list before you build on any of them.
seedance-2.5: 4 to 30 seconds, 480p/720p/1080p, native audio, and it accepts reference audio.wan-3.0: 2 to 30 seconds, native audio, 480p/720p/1080p.minimax-h3: 5 to 15 seconds, always native stereo audio.gemini-omni-flash-1.1: 3 to 10 seconds, native audio, but no audio reference.kling-3: 4 to 15 seconds at 1080p with an audio toggle.
| Model id | Length range | Audio behavior |
|---|---|---|
| seedance-2.5 | 4-30 s | Native audio, reference audio accepted |
| wan-3.0 | 2-30 s | Native audio |
| minimax-h3 | 5-15 s | Always native stereo audio |
| gemini-omni-flash-1.1 | 3-10 s | Native audio, no audio reference |
| kling-3 | 4-15 s | Audio toggle |
Which path fits which job
Pick Sume when you want a hosted API, a job id you can poll, a webhook, and per-second billing with no GPU to babysit. Pick a self-hosted Kandinsky when you need open weights, offline generation or full control of the checkpoint, and you have a card in the class the repository names.
Length is the practical difference. Kandinsky 6.0 stops at 5 seconds per clip, according to the vendor pages. Sume's Seedance 2.5 and Wan 3.0 routes accept up to 30 seconds in one request, so a single spoken beat does not have to be stitched from several generations.
If your need is an exact spoken script rather than a generated voice, that is a separate job. Sume's docs say a speaking face is a lip-sync route (Fabric with an accepted still, or MiniMax H3 Max Lip Sync) fed with audio from the TTS endpoints, because video models do not lip-sync to TTS or voice-over you supply.
How to check the catalog yourself
Ask the API what it serves today and look for the id you need. This is a read-only call and costs nothing.
import os, json, urllib.request
req = urllib.request.Request(
"https://api.sume.com/v1/videos/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
)
with urllib.request.urlopen(req) as resp:
body = json.load(resp)
print(json.dumps(body, indent=2)[:2000])What to do next
If you only need a sound-capable clip today, start from the table above and the video generation docs. If Kandinsky is the model you want, self-host it from the repository and treat Sume as the place to generate longer shots, add TTS narration, or render lip-sync from a script. The API reference covers authentication and job polling for the Sume side.
Sources
Related posts
More in Models
- Kandinsky 6.0 Video is MIT: can you use the clips commercially?
Kandinsky 6.0 Video Pro (29B) and Lite (3B) are MIT-licensed with joint audio. What MIT covers, what to check, and the hosted alternative on Sume.
- Kimi K3 on Sume: no catalog row, and what to pick for a video agent
Kimi K3 is not in Sume's agent model catalog. Moonshot says it takes text, images and video; here is the verified alternative and how Kimi can still call Sume.
- Korean clip? Whistle's seven languages vs Sume STT
Cactus Whistle lists English, German, French, Spanish, Italian, Dutch and Polish. For a Korean clip use Sume STT with a ko hint, then caption it.
- Muse Spark 1.3 on Sume: picker row, tool use and media jobs
Meta says Muse Spark 1.3 uses about 20% fewer tool calls. Sume lists it as a catalog row behind the OpenRouter switch; the API cannot pick it.
Written by Sume