Kandinsky 6.0 Lite takes 406 s on an RTX 5090: local or hosted?

The Kandinsky repo times one 5 s Full HD clip per GPU: 406 s for Lite on an RTX 5090. Clips per hour by card, and when a hosted Sume job is simpler.

5 min readSume
All posts

On an RTX 5090, Kandinsky 6.0 Lite needs 406 seconds to make one 5-second Full HD clip with sound, so a single card makes about 9 of them in an hour. On an RTX 5060 Ti the same job takes 1,774 seconds, about 2 per hour. A hosted Sume job needs no card at all, and the video docs say generation usually takes 30 seconds to several minutes depending on the model and settings.

The timings come from the performance table in the Kandinsky 6.0 repository (read 2026-10-11), which measures a 5-second video on the non-distilled model. The repository also lists a lite-distill checkpoint, but this table does not time it, so the numbers below are the slow path.

Seconds per 5-second clip, by GPU

Clips per hour is 3,600 divided by the Full HD time, rounded to one decimal place. It assumes one clip at a time and ignores load time.

Working time in seconds for a 5-second video, non-distilled model, from the Kandinsky 6.0 repository (read 2026-10-11). Clips per hour is our division.
GPULite Full HD (s)Lite clips/hourPro Full HD (s)Pro clips/hour
RTX 5060 Ti1,7742.03,5301.0
RTX 40905786.21,2472.9
RTX 50807704.71,5462.3
RTX 50904068.98544.2
RTX PRO 60003879.37654.7
A100 80 GB6645.41,1063.3
H10028412.74029.0

What the table does not tell you

The repository says it needs an NVIDIA GPU and Python 3.13 or 3.14, and that consumer cards use block offload while professional GPUs use module offload. Offload is how a big model fits a small card, and the table shows its cost: the 5060 Ti takes 3,080 seconds for Pro at SD. Treat any figure here as one measurement on the vendor's setup, not a promise for yours.

The clip is also 5 seconds with no way past that in one generation. If your product needs 8, 10 or 15 seconds, you are stitching.

The hosted side, without invented numbers

Sume's video docs describe the flow as submit, receive a job id, poll or take a webhook, then download. They give a typical range of 30 seconds to several minutes rather than a per-model guarantee, so this post does not publish a Sume-versus-GPU stopwatch. What you can compare fairly is the shape of the work:

  • Local: install, GPU drivers, checkpoint download, one clip at a time per card, and you pay for the card whether it is busy or not.
  • Hosted: one HTTP request per clip, a price per clip, and your plan's concurrency limit decides how many run at once.
  • Both: write the prompt, check the result, and budget for rejected takes.

When local wins

Local wins when the GPU already exists and sits idle at night, when you need the open MIT weights for research, or when nothing may leave the machine. Hosted wins when volume is bursty, when nobody on the team wants to maintain a CUDA stack, or when you need a clip longer than 5 seconds. Sume does not carry Kandinsky, so a hosted comparison means a different model, not the same weights; judge the output, not the label.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume