HunyuanVideo 1.5 step-distilled: 8 or 12 steps for 480p I2V
HunyuanVideo-1.5 ships a step-distilled 480p image-to-video model that needs 8 or 12 steps. What the README claims, the model sizes, and when hosted is easier.

HunyuanVideo-1.5 includes an I2V step-distilled model at 480p that, per the repository README, generates in 8 or 12 steps, with a stated 75% speedup on an RTX 4090. The base model is 8.3B parameters, and the README gives 14GB as the minimum GPU memory with offloading on. Check the license before you plan commercial use, because it is not Apache 2.0.
What is HunyuanVideo-1.5?
The README describes an 8.3B-parameter diffusion transformer with a 3D causal VAE that compresses 16 times spatially and 4 times temporally, using Selective and Sliding Tile Attention (SSTA). Default inference settings are 50 steps, 121 frames, 16:9 and a CFG scale of 6 (1 for the distilled variants).
It targets Linux with NVIDIA CUDA and Python 3.10 or newer.
Which variants exist?
The README lists variants by resolution. Pick the row that matches your task and hardware.
| Resolution | Variants listed |
|---|---|
| 480p | T2V, I2V, T2V cfg-distill, I2V cfg-distill, I2V step-distill |
| 720p | T2V, I2V, I2V cfg-distill, I2V sparse cfg-distill, SR models |
| 1080p | SR (super-resolution) model |
How fast is the step-distilled model?
The README says the 480p I2V step-distilled model generates videos in 8 or 12 steps, a 75% speedup on an RTX 4090, and that a single RTX 4090 reaches generation within 75 seconds. It also reports a 1.87 times end-to-end speedup for a 10-second 720p video against FlashAttention-3 using the sparse attention variant. These are the project's own figures; I did not reproduce them.
Fp8 GEMM inference support was added on December 23, 2025, and training code with LoRA tuning on December 5, 2025.
What license applies?
The LICENSE file in the repository is the Tencent Hunyuan Community License Agreement. It does not apply in the European Union, United Kingdom or South Korea, requires permission from Tencent if your products exceed 100 million monthly active users, and bars using outputs to improve competing AI models. It also requires you to disclose the actual provider's name to end users. Read the file yourself before shipping; I am summarizing, not giving legal advice.
When is a hosted route the easier choice?
If you are in a region the license excludes, or you cannot run Linux and CUDA, you cannot rely on the weights. Sume's docs do not list a HunyuanVideo id; run GET /v1/videos/models to confirm what exists today, and see HunyuanVideo API: hosted or weights for the longer discussion. The Video generation docs show how to pick a listed model and poll a job.
Which variant should I start with?
Start with the one that matches your input. If you have a still image and want a quick draft, the 480p I2V step-distilled model is the one the README advertises for speed. If you only have text, the 480p T2V and its cfg-distilled version are listed. Move to 720p and the SR models only after the 480p result looks right, because each step up costs time and memory.
Keep the CFG scale straight: the README sets 6 for standard models and 1 for distilled variants. Using the wrong value with a distilled model is a common reason for odd results, though I have not tested that here.
Offloading is what gets you down to the 14 GB minimum the README mentions; expect it to be slower than running fully on the GPU. Treat the 75-second figure as an upper-end result on a 4090 for the configuration the README describes, and time your own runs before you promise anyone a turnaround.
Sources
Related posts
More in Models
- Is Wan 2.7 open source? The official Wan-AI org lists Wan 2.2
Wan-AI's Hugging Face page lists Wan 2.2 weights and no 2.5, 2.6, 2.7 or 3.0 weights. Here is how to check, and which hosted route Sume lists for Wan.
- Kling 3.0 4K in Sume's Videos panel: what actually gets submitted
The panel lists 4K for Kling 3.0 but its live submit maps 4K down to 1080p. Where 4K does exist on Sume's API, and how to check the model before you pay.
- Kling 4.0 Flash makes 20-second clips: which Sume model reaches 20 s?
Kling 4.0 Flash is reported at 3-20 seconds and 720p for Ultra Yearly users. On Sume, kling-3 stops at 15 s; seedance-2.5 and wan-3.0 reach 20 s and beyond.
- Kling motion control keep_original_sound: get a silent clip
keep_original_sound on Sume's Kling 3.0 Motion Control defaults to true, so the driving video's audio rides along. Send false for a silent clip.
Written by Sume