How much VRAM for AI video generation? Numbers from model cards

From 14 GB to 80 GB: the VRAM figures Tencent, Wan-AI and Genmo publish for their open video models, and what each number was measured with.

5 min readSume
All posts

It depends on the model and the settings: the model cards read for this post state anywhere from 14 GB (HunyuanVideo-1.5 with offloading) to 80 GB (Wan2.2 A14B's single-GPU command). No one number applies to "AI video", because each card measured a different model, resolution, and memory-saving setting.

Every figure is the model authors' own, from their Hugging Face cards (HunyuanVideo-1.5, HunyuanVideo, Wan2.2 A14B, Wan2.2 TI2V-5B, Mochi 1), read 2026-09-29. The hosted-model note uses Video generation.

What VRAM does each open video model state?

From each model card on Hugging Face, read 2026-09-29. Figures are not measured on the same basis.
ModelStated figureBasis on the card
HunyuanVideo-1.514 GB minimumWith model offloading enabled
HunyuanVideo60 GB minimum; 80 GB recommended60 GB for 720px1280px129f, 45 GB for 544px960px129f
Wan2.2 T2V-A14BAt least 80 GBThe single-GPU command; offload flags reduce usage
Wan2.2 TI2V-5BAt least 24 GBThe single-GPU command, for example an RTX 4090
Mochi 1 previewAbout 60 GB; 42 GB; 22 GBGenmo's repository; Diffusers highest quality; Diffusers bfloat16

Why do the numbers differ so much?

Three things change the figure. The model size differs: the Wan2.2 card lists a 5B model beside the A14B pair, and HunyuanVideo-1.5 is described as 8.3B parameters. The settings differ: offloading moves weights to CPU memory, which is why the HunyuanVideo-1.5 card measures 14 GB with it on. And the output differs: the older HunyuanVideo card gives a lower figure for a smaller frame size than for 720px1280px.

The Mochi card is the clearest case: the same model is listed at about 60 GB in Genmo's repository, 42 GB in a Diffusers example, and 22 GB with a bfloat16 variant that the card says causes a slight drop in quality.

Which model cards give no VRAM number?

The LTX-2 card lists Python, CUDA and PyTorch versions but no minimum VRAM, and MiniMax's H3 open-source page gives a four-GPU example (--num-gpus 4) rather than a minimum. So neither is in the table. Do not fill the gap with a number from a forum post; run the model's own example and read the peak memory it reports.

What if I do not have that much VRAM?

  • Use the memory-saving flags the card names, such as Wan2.2's --offload_model True, --convert_model_dtype and --t5_cpu.
  • Pick the smaller model the same authors publish, such as Wan2.2 TI2V-5B.
  • Use a hosted model, which needs no GPU on your side.

What does a hosted video API need instead?

Sume's video generation is an asynchronous API: you send a request to POST /v1/videos, receive a job id, and poll for the result, so you need no GPU on your machine. Which models you can call is listed at GET /v1/catalog; check it rather than assuming a model from a card above is there. Do you need a GPU for AI video? covers the trade-off.

Sources

Related posts

More in Models

All Models posts

Written by Sume