Wan 2.2 TI2V-5B: 5 s of 720p in under nine minutes on a consumer GPU

Wan 2.2's card says TI2V-5B makes 5 seconds of 720P video in under nine minutes on one consumer GPU. What that means for a queue, and a hosted comparison.

4 min readSume
All posts

Wan 2.2's model card says its 5B model can generate a 5-second 720P video "in under 9 minutes on a single consumer-grade GPU". The README sets the memory floor for TI2V-5B at 24 GB (an RTX 4090 is the example). That is a ceiling on time (about 540 seconds per clip), not a typical figure, and it is one GPU doing one clip at a time.

What does 540 seconds per clip add up to?

If one card makes at most 1 clip every 9 minutes, it makes at least 6.67 clips an hour, or about 160 in 24 hours of non-stop work. Those are arithmetic from the card's ceiling; the card gives no average. Real use has loading, failed runs and idle time.

Throughput implied by the card's nine-minute ceiling (arithmetic, read 2026-10-09)
Clips wantedHours on one GPU at 9 min each
101.5
507.5
20030

How does a hosted clip compare?

Wan 2.2 is not on Sume; wan-3.0 is, as a hosted model, and it is a different model from these weights. fal lists wan-3.0 at $0.10 per second at 720p (fal, read 2026-10-09), and Sume reserves list times 1.25, so $0.125 per second, or $0.625 for 5 seconds at 720p. Read the live price from pricing_skus first.

So 200 clips would be $125 hosted, against 30 GPU-hours plus setup. Break-even is your GPU hourly rate: $125 divided by 30 is $4.17 per hour at full use. Below that and with the card sitting busy, self-hosting is cheaper per clip; above it, or with idle time, the hosted clip is.

How would you raise the throughput?

The simplest way is a second card, which doubles the ceiling and doubles the hardware bill. Memory-saving options in the README's commands can affect speed, so test with and without them.

Also consider clip length. The card's figure is for 5 seconds; a longer clip takes longer, and I found no scaling figure. If you need many 5-second takes, plan around the nine-minute ceiling; if you need long clips, hosted models such as wan-3.0 take 2 to 30 seconds per job.

What else should be on the list?

  • Quality is not compared here: different models.
  • Apache 2.0 for the weights; the hosted call has its own terms.
  • Queue latency: a hosted call returns a job and a polling_url, and the video docs say generation takes 30 seconds to several minutes.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume