Is AI video 'Full HD' native or upscaled? Kandinsky 6.0 vs Sume rows
Kandinsky 6.0 renders 864x480 and adds Full HD with a plug-in. Sume's H3 Max 1080p is a latent refinement from 768p. What that means for price and detail.

Not always. Kandinsky 6.0 Video renders 864x480 and a separate plug-in super-resolution model raises it to 1920x1080, and on Sume the minimax-h3-max 1080p tier is documented as a latent refinement from native 768p. A '1080p' label in a model list can mean a native render, a refinement, or an upscale, and the price often follows the label.
This page lays out what the two sources actually state, read on 2026-10-11, so you can tell which rows are which before you pay for the larger size.
What Kandinsky states
The kandinsky-6 repository gives the base output as 864x480 and says a plug-in super-resolution model raises the result to Full HD (1920x1080). Kandinsky Lab's site calls it built-in super-resolution up to Full HD. The two descriptions agree on the destination: a two-stage pipeline, generate small, then resolve to larger.
What Sume's docs state
Sume's video generation docs are specific about two MiniMax rows. minimax-h3 accepts 5 to 15 seconds at native 480p/768p, and 768p is first-class, not 720p. The docs add that Sume bills 2K and 4K upscales if your request includes them. minimax-h3-max is the faster 768p variant, and its 1080p is a latent refinement from native 768p. For other ids the docs list resolutions without saying how each tier is produced, so this page makes no native-or-upscaled claim about them.
| Row | Native render | Larger size | Stated as |
|---|---|---|---|
| Kandinsky 6.0 Video | 864x480 | 1920x1080 | plug-in super-resolution model |
| minimax-h3 | 480p, 768p | 2K, 4K | upscales that Sume bills when requested |
| minimax-h3-max | 768p | 1080p | latent refinement from native 768p |
What it does to the bill
On Sume the larger tier costs more per second, even when it is a refinement. These are the catalog rates, read from the code on 2026-10-11.
| Sume id and tier | Per second | 5 s clip | 10 s clip |
|---|---|---|---|
| minimax-h3, 768p | $0.075 | $0.375 | $0.75 |
| minimax-h3, 2K | $0.1625 | $0.8125 | $1.625 |
| minimax-h3, 4K | $0.20 | $1.00 | $2.00 |
| minimax-h3-max, 768p | $0.10 | $0.50 | $1.00 |
| minimax-h3-max, 1080p | $0.20 | $1.00 | $2.00 |
A workable test
A label tells you less than a frame does. To decide whether the larger tier earns its price, render the same prompt at the native tier and at the larger one, then look at fine texture such as hair, text and fabric weave. If the larger frame shows no extra real detail, the cheaper tier is your answer.
The same idea applies to Kandinsky. Compare its 864x480 output with the super-resolved Full HD output before you decide the second stage is worth its run time. The repository's timing table separates the two sizes as HD and SD rows, but I cannot tell from the page whether HD there means the 1080p plug-in stage, so ask before you budget from it.
Before you pick a tier
- Read the row's docs line, not just its resolution list.
- Draft at the native tier, then re-render the one you keep at the larger tier.
- Keep the prompt identical across tiers so any difference is the tier, not the wording.
- Pin the model id when you compare. A
sume/autorequest never names the family that ran.
Sources
Related posts
More in Comparisons
- Is Kandinsky 6.0 Video on Sume? No: query the catalog by need
Sume lists no Kandinsky 6.0 id. Map each Kandinsky feature to a Sume request field, then filter GET /v1/videos/models for sound and a 5 s duration.
- Kandinsky 6.0 lip-syncs in one pass; on Sume a talking face is Fabric
Kandinsky 6.0 Video makes speech and lip-sync inside the clip. Sume's docs send on-camera speech to Fabric or H3 Max Lip Sync with your audio, not a video id.
- Kandinsky 6.0 Pro: 292 s a clip on an H100, 100 clips take 8.1 hours
Kandinsky's published timing is 292 seconds per 5-second Pro HD clip on an H100. That is 8.1 hours for 100 clips, set against hosted per-clip prices on Sume.
- Is Kandinsky 6.0 Video on Sume? No, and the closest audio ids
Kandinsky 6.0 Video is open weights with no hosted API, and Sume lists no Kandinsky id. These Sume ids give 5-second clips with sound and a first frame.
Written by Sume