Wan 2.2 CUDA out of memory: offload_model, dtype and t5_cpu flags
The Wan 2.2 README fixes OOM with three flags: offload_model, convert_model_dtype and t5_cpu. Which commands need 80 GB and 24 GB, and what it does not promise.

For a CUDA out-of-memory error on Wan 2.2, the README's advice is to add --offload_model True, --convert_model_dtype and --t5_cpu to generate.py, which it says reduce GPU memory use. Its single-GPU A14B commands say they need a GPU with at least 80 GB of VRAM, even with those flags in the example, and its 5B text-image-to-video command says it runs on at least 24 GB, such as an RTX 4090. The README does not give a per-flag saving, so test each. Read 2026-10-08.
What the README says, by model
The same three flags appear across the single-GPU commands, with different memory floors.
| Checkpoint | Task | Resolution in command | Memory the README states |
|---|---|---|---|
| T2V-A14B | Text-to-video | 1280*720 | At least 80 GB VRAM |
| I2V-A14B | Image-to-video | 1280*720 | At least 80 GB VRAM |
| TI2V-5B | Text or image to video | 1280*704 | At least 24 GB VRAM (e.g. RTX 4090) |
The 5B command
This is the README's single-GPU text-to-video command for the 5B model, which includes all three flags:
python generate.py --task ti2v-5B --size 1280*704 \
--ckpt_dir ./Wan2.2-TI2V-5B \
--offload_model True --convert_model_dtype --t5_cpu \
--prompt "Two anthropomorphic cats in comfy boxing gear fight on a spotlighted stage"What the flags do not fix
Offloading moves weights to CPU memory, so you trade GPU memory for system RAM and speed. The README does not publish timings for the offloaded runs, so a slow first clip is not a bug report. If the A14B model still runs out of memory on a 40 GB or 48 GB card, the README's route is the multi-GPU command with FSDP and DeepSpeed Ulysses, for example torchrun --nproc_per_node=8.
The prompt-extension step, if you turn it on, loads another model, as another post on this site explains, and can bring the OOM back.
When to stop tuning
Each failed run costs time. If the goal is a handful of clips, compare the time with a hosted price. Sume lists wan-3.0, a different model from Wan 2.2, at 480p, 720p and 1080p: $0.3125, $0.625 and $1.25 for five seconds (list x 1.25). Sume does not list Wan 2.2, so a hosted run cannot reproduce a Wan 2.2 look exactly.
Order to try the flags
Add --offload_model True first, since it moves the model to CPU between steps. Add --convert_model_dtype next to change the weight precision, then --t5_cpu to keep the text encoder on the CPU. Lower the resolution only after that. Because the README gives no per-flag saving, change one flag at a time and note peak memory from nvidia-smi after each run.
If it still fails
Check which task you run. The README states 80 GB for the A14B single-GPU commands and 24 GB for the 5B model, so an A14B run on a 24 GB card is outside the stated range whatever flags you add. Switching to TI2V-5B at 1280*704 is the README's route for a 24 GB card. The alternative is the multi-GPU command with FSDP, which splits the model across cards, or a hosted job that carries none of this setup.
Do not read the 5B result as a quality match for A14B; the README presents them as separate checkpoints.
- A14B on one card: 80 GB.
- 5B on one card: 24 GB.
- More cards:
torchrunwith FSDP and Ulysses, per the README.
Sources
Related posts
More in Models
- Wan 2.2 prompt extension: a DashScope key or a local Qwen model
Wan 2.2's README offers prompt extension via Alibaba DashScope (DASH_API_KEY) or a local Qwen model. What each needs, and what Sume's request takes.
- Wan 3.0 leads text-to-video at 1,156 Elo: $12 listed, $15 on Sume
Wan 3.0 is first on the AA-Video-T2V v2.0 board at 1,156 Elo and $12.00 a minute. On Sume, wan-3.0 at 1080p bills $0.25 a second, which is $15.00 a minute.
- Wan 3.0 reference limits: 10 images, 5 videos and 5 audio tracks
Sume lists Wan 3.0 references as up to 10 images, 5 videos (15 s total) and 5 audio tracks (15 s total). Sume docs do not claim 50 references.
- What a 13-point Elo gap means: Wan 3.0 vs Seedance 2.5 win chance
Wan 3.0 leads Seedance 2.5 by 13 Elo on the AA text-to-video board. On the standard Elo scale that is a 51.9% win chance per vote. Arithmetic for other gaps.
Written by Sume