Wan 2.2 TI2V-5B in Diffusers: 121 frames, 24 fps, 24GB GPU

A working Diffusers example for Wan2.2-TI2V-5B: 704x1280 portrait frames, 121 frames at 24 fps, 24GB VRAM minimum, and when to use a hosted route instead.

5 min readSume
All posts

Wan2.2-TI2V-5B is the Wan 2.2 checkpoint that fits a single 24GB card such as an RTX 4090, and it makes 5 seconds (121 frames) at 24 fps in 720P. The model card's Diffusers example runs at 704 by 1280 with 50 steps and guidance 5.0. The card says the model needs features currently only in the main branch of diffusers, so install from GitHub.

What does the model card's example look like?

This is the example from the Wan-AI model card, trimmed to the essentials. Install diffusers from the main branch first with pip install git+https://github.com/huggingface/diffusers.

import torch
from diffusers import WanPipeline, AutoencoderKLWan
from diffusers.utils import export_to_video

model_id = "Wan-AI/Wan2.2-TI2V-5B-Diffusers"
vae = AutoencoderKLWan.from_pretrained(
    model_id, subfolder="vae", torch_dtype=torch.float32
)
pipe = WanPipeline.from_pretrained(
    model_id, vae=vae, torch_dtype=torch.bfloat16
)
pipe.to("cuda")

frames = pipe(
    prompt="Two anthropomorphic cats in comfy boxing gear fight on stage.",
    negative_prompt="low quality, blurry, static",
    height=704,
    width=1280,
    num_frames=121,
    guidance_scale=5.0,
    num_inference_steps=50,
).frames[0]

export_to_video(frames, "output.mp4", fps=24)

What are the numbers to remember?

The card gives the following settings. The VRAM line is the minimum the card states, with 80GB or more recommended.

Wan2.2-TI2V-5B settings from the model card (read 2026-10-02)
SettingValue
720P size1280x704 or 704x1280
Frame rate24 fps
Frame count121 (5 seconds)
VRAMAt least 24GB, for example an RTX 4090
LicenseApache 2.0

Why is the VAE loaded in float32?

The example loads the VAE with torch.float32 and the transformer in bfloat16. I have not measured why the card does this, so I only repeat it: copy the dtype split from the card rather than casting everything to bf16, then change one thing at a time if you need to save memory.

When would I send this prompt to a hosted route instead?

Use the local pipeline for experiments where you control the model and want to iterate for free after the hardware cost. Use a hosted route when you need a clip longer than 5 seconds, want 1080p, or are generating from a laptop. Sume's wan-3.0 row accepts 2 to 30 seconds at 480p, 720p or 1080p and bills per second by resolution; it is a newer model, so expect different output from the 5B checkpoint.

Submit with POST /v1/videos, poll the job, and download unsigned_urls[0], as in the Video generation docs. The Jobs and results page covers statuses and failures. Remember that the request has no seed field, so reruns will not be pixel-identical.

What should I check before shipping local output?

Read the license text in the repository you downloaded from, keep the model card's warning about lawful content in mind, and note that the example ships no audio. If you need a soundtrack, add it in a separate step.

What can go wrong on a first run?

Three things are worth checking before you blame the model. First, the diffusers version: the card says the pipeline needs features only in the main branch, so a release from your package index may not have them. Second, the size: the card names 1280x704 or 704x1280 for 720P, and the example uses height 704 and width 1280, so keep to those pairs rather than rounding to 720.

Third, the frame count. The example uses 121 frames, which at 24 fps is just over 5 seconds. If you change the frame count, change the fps you pass to export_to_video only if you intend to alter playback speed; the card sets 24 fps and I have not tested other values.

If you run out of memory, the card's own minimum is 24GB and its recommendation is 80GB or more, so a smaller card is outside what the card states. Rather than hunting for undocumented tricks, drop to a shorter run or test the same idea on a hosted id first.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume