HunyuanVideo 1.5 low VRAM: 14 GB with offloading

Tencent's HunyuanVideo-1.5 card lists 14 GB as the minimum GPU memory with offloading on. The OOM settings and step-distilled option it names.

4 min readSume
All posts

Tencent's HunyuanVideo-1.5 model card lists a minimum of 14 GB of GPU memory, measured with model offloading enabled, on an NVIDIA GPU with CUDA and Linux. If you still hit out-of-memory errors, the card gives one setting to try for GPU memory and one for CPU memory.

Everything here is from the HunyuanVideo-1.5 card on Hugging Face, read 2026-09-29. The hosted-model note at the end uses Video generation.

What are the hardware requirements?

The card's System Requirements section is short, and the memory line carries a caveat: the figure was measured with offloading enabled. With enough GPU memory you can turn offloading off for faster inference.

From the HunyuanVideo-1.5 model card, read 2026-09-29.
RequirementWhat the card says
GPUNVIDIA GPU with CUDA support
Minimum GPU memory14 GB (with model offloading enabled)
Operating systemLinux
PythonPython 3.10 or higher

What do I do if I still get out-of-memory errors?

The card gives one tip for GPU memory and one for CPU memory:

  • GPU memory above 14 GB but still out of memory: set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:128 before running.
  • Limited CPU memory: turn off overlapped group offloading with --overlap_group_offloading false.
  • The card describes overlapped group offloading as enabled by default; it "significantly increases CPU memory usage but speeds up inference", so it is a trade between CPU RAM and speed.

Is there a faster option on a consumer GPU?

The card's Dec 05, 2025 news item says a 480p image-to-video step-distilled model generates videos in 8 or 12 steps (recommended), and that on an RTX 4090 a single card can generate videos within 75 seconds. You enable it with the --enable_step_distill parameter. The card also notes FP8 GEMM inference support from Dec 23, 2025.

These are the authors' numbers for their own setup and the 480p image-to-video variant, so do not read them as a promise for text-to-video or 720p.

What tools can run it?

The card's news list says HunyuanVideo-1.5 is available in Hugging Face Diffusers, and it links a ComfyUI usage guide. It also lists cache inference (deepcache, teacache and taylorcache, from Nov 27, 2025) as a way to speed up generation, and step-distilled inference as a separate option.

Under community contributions the card lists Wan2GP and describes it as "a very low VRAM app (as low 6 GB of VRAM for Hunyuan Video 1.5)". That is a third-party app the card links, not a figure Tencent measured, so treat it as a lead to check.

Can I skip the GPU and call a hosted model?

Not for this model on Sume: Sume's docs and code list no HunyuanVideo id. If you only need a finished clip, the models the catalog does list run on the provider's side; GET /v1/catalog is the place to check, as described on Video generation.

For the wider question of when to run weights yourself, see Open-source video model vs API.

Sources

Related posts

More in Models

All Models posts

Written by Sume