HunyuanVideo 1.5 low VRAM: 14 GB with offloading
Tencent's HunyuanVideo-1.5 card lists 14 GB as the minimum GPU memory with offloading on. The OOM settings and step-distilled option it names.

Tencent's HunyuanVideo-1.5 model card lists a minimum of 14 GB of GPU memory, measured with model offloading enabled, on an NVIDIA GPU with CUDA and Linux. If you still hit out-of-memory errors, the card gives one setting to try for GPU memory and one for CPU memory.
Everything here is from the HunyuanVideo-1.5 card on Hugging Face, read 2026-09-29. The hosted-model note at the end uses Video generation.
What are the hardware requirements?
The card's System Requirements section is short, and the memory line carries a caveat: the figure was measured with offloading enabled. With enough GPU memory you can turn offloading off for faster inference.
| Requirement | What the card says |
|---|---|
| GPU | NVIDIA GPU with CUDA support |
| Minimum GPU memory | 14 GB (with model offloading enabled) |
| Operating system | Linux |
| Python | Python 3.10 or higher |
What do I do if I still get out-of-memory errors?
The card gives one tip for GPU memory and one for CPU memory:
- GPU memory above 14 GB but still out of memory: set
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:128before running. - Limited CPU memory: turn off overlapped group offloading with
--overlap_group_offloading false. - The card describes overlapped group offloading as enabled by default; it "significantly increases CPU memory usage but speeds up inference", so it is a trade between CPU RAM and speed.
Is there a faster option on a consumer GPU?
The card's Dec 05, 2025 news item says a 480p image-to-video step-distilled model generates videos in 8 or 12 steps (recommended), and that on an RTX 4090 a single card can generate videos within 75 seconds. You enable it with the --enable_step_distill parameter. The card also notes FP8 GEMM inference support from Dec 23, 2025.
These are the authors' numbers for their own setup and the 480p image-to-video variant, so do not read them as a promise for text-to-video or 720p.
What tools can run it?
The card's news list says HunyuanVideo-1.5 is available in Hugging Face Diffusers, and it links a ComfyUI usage guide. It also lists cache inference (deepcache, teacache and taylorcache, from Nov 27, 2025) as a way to speed up generation, and step-distilled inference as a separate option.
Under community contributions the card lists Wan2GP and describes it as "a very low VRAM app (as low 6 GB of VRAM for Hunyuan Video 1.5)". That is a third-party app the card links, not a figure Tencent measured, so treat it as a lead to check.
Can I skip the GPU and call a hosted model?
Not for this model on Sume: Sume's docs and code list no HunyuanVideo id. If you only need a finished clip, the models the catalog does list run on the provider's side; GET /v1/catalog is the place to check, as described on Video generation.
For the wider question of when to run weights yourself, see Open-source video model vs API.
Sources
Related posts
More in Models
- HunyuanVideo API: is there a hosted one, or only weights?
HunyuanVideo is published as downloadable weights. Sume's docs list no HunyuanVideo API, so here is what running it takes and the listed video ids.
- Ideogram 4.0 API: request shape and what Sume lists
Ideogram 4.0 has its own endpoint with text_prompt or json_prompt. Sume's image catalog lists Ideogram V3 only; here is how the two differ.
- Imagen 4 API shut down in the Gemini API: what Sume's catalog lists
Google says Imagen is shut down in the Gemini API and points to Nano Banana. Sume's image catalog lists google/imagen-4-fast, imagen-4-ultra and Nano Banana ids
- Kling 2.6 vs 3.0: length, resolution, audio and price
Kling 3.0 makes 3–15 s clips up to 4K with audio and multi-shot; 2.6 makes 5 or 10 s shots at 720p or 1080p. Kling's own specs and prices, dated.
Written by Sume