Is Kandinsky 6.0 Video on Sume? No: query the catalog by need
Sume lists no Kandinsky 6.0 id. Map each Kandinsky feature to a Sume request field, then filter GET /v1/videos/models for sound and a 5 s duration.

No: Kandinsky 6.0 Video is not in Sume's video catalog, and nothing in the API source or docs of the Sume repository names it. Sume lists hosted ids only, so open-weights releases like this one are not callable through POST /v1/videos. What you can do is describe what you wanted from Kandinsky (sound, 5 seconds, a first frame) and filter the live catalog for ids that match.
The Kandinsky facts are from the project repository and the technical report, read 2026-10-11. Sume's request fields are from Video Generation. The list of ids changes, so treat the script at the end as the source of truth and this page as the map.
How this was checked
I searched the Sume checkout this post was written against (docs, API source, web app and packages) for the name Kandinsky and found no hit. The docs name the discovery route directly: GET /v1/videos/models returns every video id with its supported_durations, supported_resolutions, generate_audio and pricing_skus. That endpoint, not a blog table, decides what you can send today.
Kandinsky feature to Sume field
Each thing you would do with Kandinsky has a place on the Sume request, if a listed model supports it.
| If you wanted from Kandinsky | Sume request or catalog field | Check before sending |
|---|---|---|
| A 5-second clip | duration: 5 | supported_durations includes 5 |
| Synchronized audio | generate_audio, default is the model's own capability | generate_audio is true in the catalog |
| Image-to-audio-video | frame_images with frame_type: first_frame | supported_frame_images |
| Full HD output | resolution: 1080p | supported_resolutions |
| A fixed seed for a rerun | Not available: every v1 model reports seed: false | None; a seed is rejected |
| Provider-specific options | Not available: provider.options must be empty | allowed_passthrough_parameters is empty |
What differs in practice
Kandinsky's repository lists inference defaults for Pro of 50 steps and guidance 5.0. Those are knobs on your machine. Sume exposes none of them: the portable controls are prompt, duration, resolution, aspect ratio, first or last frame, reference media and the generate_audio flag. If your workflow depends on step count or guidance, that is a reason to run open weights yourself.
The other difference is lifecycle. A Sume video request is asynchronous. You get a job id, poll the polling URL or pass an HTTPS callback_url, then download from unsigned_urls. Send an Idempotency-Key so a retry does not create a second job.
Script: which listed ids fit a Kandinsky-shaped job
Run this with SUME_API_KEY set. It prints how many ids are listed, whether any id contains the word Kandinsky, and every id that can make audio and accepts a 5-second clip.
import os
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
r = requests.get("https://api.sume.com/v1/videos/models", headers=H, timeout=30)
r.raise_for_status()
ids = [m["id"] for m in r.json()["data"]]
print(len(ids), "video ids")
print("kandinsky listed:", any("kandinsky" in i.lower() for i in ids))
for m in r.json()["data"]:
if m.get("generate_audio") and 5 in (m.get("supported_durations") or []):
print(m["id"], m["supported_resolutions"])If the list is empty or changes
If no id passes the filter, relax one constraint at a time. Drop the audio filter and add sound later with the audio tools, or allow 6 to 10 seconds and trim. Do not assume a model has a feature because a neighbor does; the docs call out, for example, that Gemini Omni Flash 1.1 takes 3 to 10 seconds, not the 15 seconds that other catalog ids top out at.
Sources
Related posts
More in Comparisons
- Kandinsky 6.0 lip-syncs in one pass; on Sume a talking face is Fabric
Kandinsky 6.0 Video makes speech and lip-sync inside the clip. Sume's docs send on-camera speech to Fabric or H3 Max Lip Sync with your audio, not a video id.
- Kandinsky 6.0 Pro: 292 s a clip on an H100, 100 clips take 8.1 hours
Kandinsky's published timing is 292 seconds per 5-second Pro HD clip on an H100. That is 8.1 hours for 100 clips, set against hosted per-clip prices on Sume.
- Is Kandinsky 6.0 Video on Sume? No, and the closest audio ids
Kandinsky 6.0 Video is open weights with no hosted API, and Sume lists no Kandinsky id. These Sume ids give 5-second clips with sound and a first frame.
- Kandinsky 6.0 image-to-audio-video vs a first frame on Sume
Kandinsky 6.0 animates a still with synced audio. On Sume, frame_images starts a clip from your image, and some ids also take an audio reference. What fits.
Written by Sume