HunyuanVideo 1.5 on 14 GB of VRAM: run locally or call an API
The HunyuanVideo 1.5 repo lists 8.3B parameters, 480p to 1080p and a 14 GB VRAM minimum with offloading. A local-run versus API checklist.

The HunyuanVideo 1.5 repository, read 2026-10-03, lists 8.3 billion parameters, output at 480p, 720p and 1080p, and a minimum of 14 GB of VRAM when offloading is on. A step-distilled variant runs in 8 to 12 steps. The repo material dates from December 2025, so it is older than the other items in this series.
The specs
These come from the repository page.
| Item | Listed |
|---|---|
| Parameters | 8.3B |
| Resolutions | 480p, 720p, 1080p |
| Minimum VRAM | 14 GB with offloading |
| Step-distilled variant | 8-12 steps |
| Source date | December 2025 |
What 14 GB means in practice
14 GB is a floor with offloading enabled, which trades memory for time by moving parts of the model in and out of GPU memory. The repo figure does not tell you how long a clip takes, so measure that on your own card before you plan around it.
A local versus API checklist
Answer these before you pick.
- Do you need to keep inputs on your own hardware? That favors local.
- Do you need many clips in parallel? An API job queue handles concurrency; on Sume, valid paid jobs are accepted as
queuedwhile queue capacity remains. - Do you need audio? Check the model card. Sume reports an audio flag per catalog model.
- Who handles failures and retries? Locally, you. On Sume, you send an
Idempotency-Keyand poll the stored job id. - What is the cost per finished clip? Local cost is your hardware and time; hosted cost is the per-second rate. Sume bills provider list times 1.25.
Measure before you plan
Run a five-clip test on your own hardware and record three numbers: seconds to finish, peak VRAM, and the number of clips you would keep. Divide your GPU cost per hour by clips kept per hour to get a cost per finished clip, then compare it with a hosted per-second rate times your clip length.
Date the test. Drivers, offloading settings and repo updates change the result.
A hybrid path
A common pattern is to use the local model for drafts and a hosted model for finals, or the reverse. Keep one small wrapper in your code that takes a prompt and returns a file, and swap the implementation behind it. With Sume, the wrapper is a submit to POST /v1/videos, a poll on the returned polling_url, and a download from unsigned_urls[0], as described in the video generation guide.
Sources
Related posts
More in Models
- Luma's 2026 timeline: Ray3.14, Ray3.2, Scenes and Variants
Luma shipped Ray3.14 in January, Ray3.2 in June, Scenes in August and Variants on Oct 1, 2026. What each added, and why to pin model ids.
- Lyria 3.5 blocks artist-voice prompts: how to write briefs that pass
Google's Lyria 3.5 docs note that prompts asking for specific artist voices are blocked. Describe the sound instead, then run it through the Sume Music Router.
- MiniMax H3 limits: 9 images, 3 videos, 3 audio, file caps
MiniMax's H3 guide caps prompts at 7,000 characters and references at 9 images, 3 videos and 3 audio files. Cheat sheet with the Sume limits beside it.
- Nano Banana reference image limits: Lite, 2 and Pro compared
Nano Banana 2 Lite takes up to 14 object images, Nano Banana 2 takes 10 object, 4 character and 3 style, Pro takes 6 object and 5 character.
Written by Sume