Wan 2.2 A14B explained: 27B total, 14B active, 80GB GPU

Wan 2.2's A14B models use two experts, one for noisy early steps and one for late detail. What that means for GPU memory, and when a hosted route is simpler.

5 min readSume
All posts

Wan 2.2's A14B models (T2V-A14B and I2V-A14B) are a two-expert mixture-of-experts design: roughly 14B parameters per expert, about 27B in total, with only one 14B expert running at each denoising step. The README still lists an 80GB single GPU for them, so "only 14B active" lowers compute per step, not the weights you must be able to hold or offload.

What are the two experts?

The Wan2.2 README describes a two-expert design tailored to the denoising process. A high-noise expert handles the early steps, where the layout and motion are decided; a low-noise expert handles the later steps, where detail is refined. The switch between them is based on a signal-to-noise threshold.

Because the expert is chosen per step, a single generation touches both, one after the other. Total parameters (about 27B) set how much has to live in memory or be swapped; active parameters (14B) set the compute per step.

What hardware does each Wan 2.2 model list?

Numbers below come from the README table, not from my own runs.

Wan 2.2 model tasks and listed hardware (read 2026-10-02)
ModelTaskResolutionListed GPU memory
T2V-A14BText-to-video480P and 720P80GB single GPU
I2V-A14BImage-to-video480P and 720P80GB single GPU
TI2V-5BText and image to video720P at 24fps24GB (RTX 4090)
S2V-14BSpeech-to-video480P and 720P80GB single GPU
Animate-14BCharacter animation or replacementNot specifiedNot specified

Does the MoE design make Wan 2.2 cheaper to run?

It makes each step cheaper than a dense 27B model would be, which is the point of the design. It does not make the model fit a gaming card. If you only have 24GB, the 5B hybrid (TI2V-5B) is the checkpoint the README points at; see the TI2V-5B example for working code.

Memory-saving flags exist in the Wan repository, but I did not measure what quality or speed they cost, so I do not quote a smaller figure.

When is a hosted route simpler than running A14B?

A 5-second clip is the natural output of the Wan 2.2 release (121 frames at 24fps in the TI2V-5B card). If your deliverable is longer or higher resolution, the arithmetic changes. Sume's wan-3.0 catalog row accepts 2 to 30 seconds and up to 1080p, and you pay per second by resolution rather than by GPU hour.

That is a different, newer model from the Wan 2.2 weights, so it is a substitute, not a mirror. Read the live limits with GET /v1/videos/models as described in the Video generation docs.

What license covers the weights?

The README names Apache 2.0, which permits commercial use. The TI2V-5B model card adds that you are responsible for making sure your content complies with applicable laws and the restrictions in the full license text. Read the LICENSE file in the repository you download from.

How should I plan a first test?

Begin with TI2V-5B if you have a 24GB card, because it is the only Wan 2.2 text-and-image model with that listed minimum. Move to A14B only when you have an 80GB card available, and compare the same prompt on both. Then run the prompt once on a hosted id to see what a newer generation does with it. Three results side by side tell you more than any spec table.

Sources

Related posts

More in Models

All Models posts

Written by Sume