Best image-to-video model, October 2026: Wan 3.0 is 7th at 1,164 Elo
On the AA image-to-video board MiniMax H3 Max leads at 1,195 and Wan 3.0 is seventh at 1,164. Which of the top 12 you can call on Sume, with per-minute prices.

On Artificial Analysis's image-to-video board (AA-Video-I2V v1.0) MiniMax H3 Max leads at 1,195 Elo and $4.80 per minute, and Wan 3.0, first on text-to-video, is seventh at 1,164. If you animate a still, the model to try first is H3 Max (minimax-h3-max on Sume, $6.00 per minute at 768p), and the second is Gemini Omni Flash 1.1 (gemini-omni-flash-1.1).
Six of the twelve entries have a Sume id. This page lists which ones, with the Sume price per minute.
Top 12 and where each runs
The board ranks animating a source image into video with audio, by Elo from pairwise human votes. The last column checks each name against Sume's video id list (twelve ids in the repo this week).
| Rank | Model | Elo | Listed per minute | Sume id |
|---|---|---|---|---|
| 1 | MiniMax H3 Max | 1195 | $4.80 | minimax-h3-max |
| 2 | MiniMax H3 | 1181 | $7.80 | minimax-h3 |
| 3 | Vidu Q4 Preview | 1179 | $7.20 | Not listed |
| 4 | Gemini Omni Flash (no version shown) | 1178 | $6.00 | gemini-omni-flash-1.1 |
| 5 | Dreamina Seedance 2.0 720p | 1176 | $9.07 | seedance-2 |
| 6 | HiDream-O1-Video-1.0 | 1175 | $5.80 | Not listed |
| 7 | Wan 3.0 | 1164 | $12.00 | wan-3.0 |
| 8 | HappyHorse-1.1 | 1104 | $9.90 | Not listed |
| 9 | Grok Imagine Video 1.5 | 1098 | $8.40 | grok-imagine-video-1.5 |
| 10 | MAGI-2 Preview | 1093 | Coming soon | Not listed |
| 11 | HappyHorse-1.0 | 1083 | $13.20 | Not listed |
| 12 | Veo 3.1 | 1082 | $24.00 | Not listed |
Sume price for the callable ones
Sume bills list times 1.25. These per-minute figures use the per-second list in Sume's price table at the resolution named.
| Sume id | Resolution | Duration range | Sume per second | Sume per minute |
|---|---|---|---|---|
| minimax-h3-max | 768p | 5 to 15 s | $0.10 | $6.00 |
| minimax-h3 | 768p | 5 to 15 s | $0.075 | $4.50 |
| gemini-omni-flash-1.1 | 720p | 3 to 10 s | $0.125 | $7.50 |
| wan-3.0 | 720p | 2 to 30 s | $0.125 | $7.50 |
| seedance-2 | 720p | 4 to 15 s | $0.378 | $22.68 |
How to send the still
Every one of these takes the still as a first frame. On /v1/videos you send it in frame_images with frame_type set to first_frame and a public HTTPS URL. Read supported_frame_images from GET /v1/videos/models to see which ids also take a last_frame. Grok Imagine Video 1.5 is image-to-video only on Sume, so it has no text-only mode.
Pick by what the still needs
Choose H3 Max for the price per minute and for 5 to 15 second clips. Choose Wan 3.0 when you need more than 15 seconds from one still, since it takes up to 30. Choose Omni when you need 1080p or 4K and a 3 to 10 second clip. Before you commit to a batch, see how the Wan 3.0 text-to-video lead compares, because the same model ranks differently by task.
Cost of a still-to-clip test
To test the board's order on your own stills, send the same first frame to the cheapest board models and one expensive one. For a 5 second clip at 720p, Wan 3.0 costs $0.625, Gemini Omni Flash 1.1 $0.625, MiniMax H3 Max at 768p $0.50 and Seedance 2.0 $1.89. The four together are $3.64.
Use the same prompt and the same still for every model and look at the first second of each clip. Motion that drifts away from the still is the usual failure, and a board score will not tell you how a given product photo behaves.
Sources
Related posts
More in Models
- Can I show HunyuanVideo clips to EU viewers? Section 5(c) read
Tencent's HunyuanVideo license says do not use or display Outputs outside the Territory, which excludes the EU, UK and South Korea. Clause by clause.
- Can you train a model on open video weights' outputs? Four licenses
H3 and HunyuanVideo bar using outputs to improve other models. LTX bars commercial distillation. Wan 2.2 is Apache 2.0. The clauses, quoted from vendor files.
- Chinese and English text in images: Hy Image 3.5's claim, Sume's rows
OpenRouter says Hy Image 3.5 Preview renders Chinese and English text. Sume does not list it. How to test Chinese signage on the Sume rows that are listed.
- Clef context: 65,536 on Cloudflare, 16,384 on the model card?
Cloudflare's Workers AI page lists 65,536 tokens of context for Clef; its Hugging Face card says 16,384 by default. Budget from the smaller. Cost math inside.
Written by Sume