HunyuanVideo 1.5 vs Wan 2.2: memory, resolution and license lines
HunyuanVideo 1.5 vs Wan 2.2 from the vendors' own pages: 14 GB with offloading vs 24 GB and 80 GB minimums, resolutions, and the license lines that differ.

HunyuanVideo 1.5 has the lower memory floor on paper: 14 GB with model offloading, for an 8.3B-parameter model. Wan 2.2 needs at least 24 GB for its TI2V-5B model and at least 80 GB for the A14B and S2V-14B models. They also differ on license: Wan 2.2's README says Apache 2.0, while HunyuanVideo 1.5's card carries the tag tencent-hunyuan-community, a separate license file you have to read.
What do the vendors say, side by side?
| Item | HunyuanVideo 1.5 | Wan 2.2 |
|---|---|---|
| Sizes named | 8.3B parameters | A14B (two experts), TI2V-5B, S2V-14B, Animate-14B |
| Minimum GPU memory | 14 GB with offloading | 24 GB (TI2V-5B); 80 GB (A14B, S2V-14B) |
| Resolutions | 480p, 720p, plus 720p to 1080p super-resolution | 480P and 720P; TI2V-5B at 720P, 24 fps |
| License | tencent-hunyuan-community tag | Apache 2.0 |
| OS note | Linux, Python 3.10 or higher | not stated in the text I read |
| Weights first public | Nov 20, 2025 | Jul 28, 2025 |
Is one of them better?
Neither page settles that. The HunyuanVideo 1.5 materials describe a GSB evaluation against other open-source systems on 300 test cases and do not name Wan 2.2 in the README I read. The Wan 2.2 card describes more training data than Wan 2.1 (65.6 percent more images, 83.2 percent more videos) and says nothing about HunyuanVideo. Run your own prompts on both.
- Memory floor is a minimum, not a speed claim.
- Check the license file for what you will do with the output, including territory and user-count lines.
- Pick by license first if you sell clips.
What should you test on your own?
Build a ten-prompt set that matches your product, and run it on both models at 720p. Time each run on your card and note memory. For HunyuanVideo, record whether prompt rewrite was on. For Wan, record which model (A14B or TI2V-5B) and whether you used offloading, because that changes both time and memory.
Score the outputs blind on the things you care about, such as motion, faces or text. Then read both licenses with your business in mind: where your users are, how many there are, and whether you train on the outputs.
Where does a hosted API fit?
Sume lists neither (catalog code read 2026-10-09). It lists wan-3.0 (2 to 30 seconds), a hosted model that is not the Wan 2.2 weights, plus minimax-h3, seedance-2.5 and others. If your comparison is about a deadline rather than weights, start with the video docs and read pricing_skus per id.
Sources
Related posts
More in Comparisons
- Hy Image 3.0 or 3.5 Preview: 3 vs 20 references, size rules, seed
Tencent's Hy Image 3.0 and 3.5 Preview side by side: reference counts, prompt limits, size ranges, seed ranges and resize_max_pixels, with what Sume lists.
- Keyterm prompting costs $0.05/hour at two vendors: 60 hours is $3.00
ElevenLabs and AssemblyAI both list keyterm prompting at $0.05 an hour: $3.00 for 60 hours of calls. Sume STT has a language hint but no keyterm list.
- Kling 3 Pro text-to-video is $0.14/s on fal; four Sume ids cost less
Sume lists no Kling text-to-video, only Kling 3.0 Motion Control at $0.1575. Four Sume text-to-video ids cost $0.075 to $0.125 per second against fal's $0.14.
- LTX-2.5 vs MiniMax H3: license lines and run requirements
LTX-2.5 vs MiniMax H3 from vendor pages: 10M vs 20M USD revenue lines, excluded territories, Python and CUDA needs, frame rules, audio, and what Sume lists.
Written by Sume