HunyuanVideo 1.5 step-distilled: 8 or 12 steps for 480p I2V

HunyuanVideo-1.5 ships a step-distilled 480p image-to-video model that needs 8 or 12 steps. What the README claims, the model sizes, and when hosted is easier.

4 min readSume
All posts

HunyuanVideo-1.5 includes an I2V step-distilled model at 480p that, per the repository README, generates in 8 or 12 steps, with a stated 75% speedup on an RTX 4090. The base model is 8.3B parameters, and the README gives 14GB as the minimum GPU memory with offloading on. Check the license before you plan commercial use, because it is not Apache 2.0.

What is HunyuanVideo-1.5?

The README describes an 8.3B-parameter diffusion transformer with a 3D causal VAE that compresses 16 times spatially and 4 times temporally, using Selective and Sliding Tile Attention (SSTA). Default inference settings are 50 steps, 121 frames, 16:9 and a CFG scale of 6 (1 for the distilled variants).

It targets Linux with NVIDIA CUDA and Python 3.10 or newer.

Which variants exist?

The README lists variants by resolution. Pick the row that matches your task and hardware.

HunyuanVideo-1.5 variants listed in the README (read 2026-10-02)
ResolutionVariants listed
480pT2V, I2V, T2V cfg-distill, I2V cfg-distill, I2V step-distill
720pT2V, I2V, I2V cfg-distill, I2V sparse cfg-distill, SR models
1080pSR (super-resolution) model

How fast is the step-distilled model?

The README says the 480p I2V step-distilled model generates videos in 8 or 12 steps, a 75% speedup on an RTX 4090, and that a single RTX 4090 reaches generation within 75 seconds. It also reports a 1.87 times end-to-end speedup for a 10-second 720p video against FlashAttention-3 using the sparse attention variant. These are the project's own figures; I did not reproduce them.

Fp8 GEMM inference support was added on December 23, 2025, and training code with LoRA tuning on December 5, 2025.

What license applies?

The LICENSE file in the repository is the Tencent Hunyuan Community License Agreement. It does not apply in the European Union, United Kingdom or South Korea, requires permission from Tencent if your products exceed 100 million monthly active users, and bars using outputs to improve competing AI models. It also requires you to disclose the actual provider's name to end users. Read the file yourself before shipping; I am summarizing, not giving legal advice.

When is a hosted route the easier choice?

If you are in a region the license excludes, or you cannot run Linux and CUDA, you cannot rely on the weights. Sume's docs do not list a HunyuanVideo id; run GET /v1/videos/models to confirm what exists today, and see HunyuanVideo API: hosted or weights for the longer discussion. The Video generation docs show how to pick a listed model and poll a job.

Which variant should I start with?

Start with the one that matches your input. If you have a still image and want a quick draft, the 480p I2V step-distilled model is the one the README advertises for speed. If you only have text, the 480p T2V and its cfg-distilled version are listed. Move to 720p and the SR models only after the 480p result looks right, because each step up costs time and memory.

Keep the CFG scale straight: the README sets 6 for standard models and 1 for distilled variants. Using the wrong value with a distilled model is a common reason for odd results, though I have not tested that here.

Offloading is what gets you down to the 14 GB minimum the README mentions; expect it to be slower than running fully on the GPU. Treat the 75-second figure as an upper-end result on a 4090 for the configuration the README describes, and time your own runs before you promise anyone a turnaround.

Sources

Related posts

More in Models

All Models posts

Written by Sume