Wan 2.2 CUDA out of memory: offload_model, dtype and t5_cpu flags

The Wan 2.2 README fixes OOM with three flags: offload_model, convert_model_dtype and t5_cpu. Which commands need 80 GB and 24 GB, and what it does not promise.

5 min readSume
All posts

For a CUDA out-of-memory error on Wan 2.2, the README's advice is to add --offload_model True, --convert_model_dtype and --t5_cpu to generate.py, which it says reduce GPU memory use. Its single-GPU A14B commands say they need a GPU with at least 80 GB of VRAM, even with those flags in the example, and its 5B text-image-to-video command says it runs on at least 24 GB, such as an RTX 4090. The README does not give a per-flag saving, so test each. Read 2026-10-08.

What the README says, by model

The same three flags appear across the single-GPU commands, with different memory floors.

Wan 2.2 single-GPU memory statements, README read 2026-10-08
CheckpointTaskResolution in commandMemory the README states
T2V-A14BText-to-video1280*720At least 80 GB VRAM
I2V-A14BImage-to-video1280*720At least 80 GB VRAM
TI2V-5BText or image to video1280*704At least 24 GB VRAM (e.g. RTX 4090)

The 5B command

This is the README's single-GPU text-to-video command for the 5B model, which includes all three flags:

python generate.py --task ti2v-5B --size 1280*704 \
  --ckpt_dir ./Wan2.2-TI2V-5B \
  --offload_model True --convert_model_dtype --t5_cpu \
  --prompt "Two anthropomorphic cats in comfy boxing gear fight on a spotlighted stage"

What the flags do not fix

Offloading moves weights to CPU memory, so you trade GPU memory for system RAM and speed. The README does not publish timings for the offloaded runs, so a slow first clip is not a bug report. If the A14B model still runs out of memory on a 40 GB or 48 GB card, the README's route is the multi-GPU command with FSDP and DeepSpeed Ulysses, for example torchrun --nproc_per_node=8.

The prompt-extension step, if you turn it on, loads another model, as another post on this site explains, and can bring the OOM back.

When to stop tuning

Each failed run costs time. If the goal is a handful of clips, compare the time with a hosted price. Sume lists wan-3.0, a different model from Wan 2.2, at 480p, 720p and 1080p: $0.3125, $0.625 and $1.25 for five seconds (list x 1.25). Sume does not list Wan 2.2, so a hosted run cannot reproduce a Wan 2.2 look exactly.

Order to try the flags

Add --offload_model True first, since it moves the model to CPU between steps. Add --convert_model_dtype next to change the weight precision, then --t5_cpu to keep the text encoder on the CPU. Lower the resolution only after that. Because the README gives no per-flag saving, change one flag at a time and note peak memory from nvidia-smi after each run.

If it still fails

Check which task you run. The README states 80 GB for the A14B single-GPU commands and 24 GB for the 5B model, so an A14B run on a 24 GB card is outside the stated range whatever flags you add. Switching to TI2V-5B at 1280*704 is the README's route for a 24 GB card. The alternative is the multi-GPU command with FSDP, which splits the model across cards, or a hosted job that carries none of this setup.

Do not read the 5B result as a quality match for A14B; the README presents them as separate checkpoints.

  • A14B on one card: 80 GB.
  • 5B on one card: 24 GB.
  • More cards: torchrun with FSDP and Ulysses, per the README.

Sources

Related posts

More in Models

All Models posts

Written by Sume