HunyuanVideo 1.5: which checkpoint to download (480p, 720p, SR)
HunyuanVideo 1.5's README names 13 checkpoints, 11 with download links: 480p, 720p, CFG- and step-distilled, sparse and SR. How to choose, with the 14 GB floor.

HunyuanVideo 1.5's README lists 13 model names; at the read date 11 have download links and 2 (720P T2V cfg-distill and 720P T2V sparse-cfg-distill) say coming soon. For a first run, choose by the job: a 480p image-to-video draft starts with the step-distilled file (8 or 12 steps), 720p work with a CFG-distilled or sparse I2V variant, and 1080p comes from a super-resolution pass. The minimum GPU memory is 14 GB, and only with model offloading enabled.
What are the checkpoints?
| Variant | What the README calls it |
|---|---|
| 480P-T2V | text-to-video |
| 480P-I2V | image-to-video |
| 480P-T2V-cfg-distill | distilled T2V |
| 480P-I2V-cfg-distill | distilled I2V |
| 480P-I2V-step-distill | step-distilled I2V, 8 or 12 steps recommended |
| 720P-T2V | text-to-video |
| 720P-I2V | image-to-video |
| 720P-I2V-cfg-distill | distilled I2V |
| 720P-I2V-sparse-cfg-distill | sparse distilled I2V |
| 720P-T2V-cfg-distill | coming soon |
| 720P-T2V-sparse-cfg-distill | coming soon |
| 720P-sr-step-distill | super-resolution |
| 1080P-sr-step-distill | super-resolution |
How should you choose?
Start from the input you have and the size you need. The README states the default is 50 inference steps and that the best configuration varies by variant (8 to 50). The step-distilled 480p I2V model is the one with a published speed claim: about 75 percent faster end to end on an RTX 4090, within 75 seconds. I found no such claim for the other files, so time them on your card.
- No input image: pick a T2V file.
- Have a first frame: pick an I2V file at the same resolution.
- Need speed: cfg-distill, then step-distill at 480p.
- Need 1080p: run the super-resolution model after a 720p generation.
What about disk and download order?
Download one file first, not every variant. A 480p I2V step-distilled checkpoint is the quickest way to prove your install works, because it runs in 8 or 12 steps. Add the 720p model once the pipeline is stable, and the super-resolution model only if you need 1080p. The README lists the variants but not their sizes, so check each file's size on the model page before you fetch it.
Keep a note of which variant made which clip. The names carry the resolution and the distillation, which is enough to reproduce a run.
Is it on Sume?
No. The Sume catalog code (read 2026-10-09) has no HunyuanVideo row. For hosted image-to-video, frame_images takes a first or last frame; the listed alternatives and their limits are on GET /v1/videos/models and in the video docs. The model card's license tag is tencent-hunyuan-community; read the license file before shipping output.
Sources
Related posts
More in Models
- Hy Image 3.5 Preview: 17.44 s median latency vs Sume's 30 s wait
OpenRouter shows a 17.44 s median for Hy Image 3.5 Preview. Sume lists no Hy row, but its image calls wait 30 s, then return a 202 job. Handle both.
- Hy Image 3.5 Preview at 91.57% availability: retry math for 200 images
OpenRouter shows 91.57% availability for Hy Image 3.5 Preview. For 200 images that is about 17 failed first tries. The math, and Sume's failed-job billing.
- Hy Image 3.5 Preview footnote watermark: 16 characters, none on Sume
Tencent's Hy Image 3.5 Preview can print a custom footnote of up to 16 characters in the lower right. Sume has no watermark field; how to add one yourself.
- Hy Image 3.5 Preview use_search_tool is off by default: what Sume has
Tencent's Hy Image 3.5 Preview has a use_search_tool switch, off by default. It is not on Sume; Nano Banana 2.1 grounding is also absent from the Sume request.
Written by Sume