How to choose an AI video model: five questions before you pin one
Length, resolution, audio, start image and references decide the model, not the leaderboard. Five questions mapped to Sume's video catalog.

Choose an AI video model by five questions, in this order: how long, how sharp, with or without sound, from what input, and with which references. Each one removes rows from Sume's catalog, and the one or two left are the ones worth testing. New releases such as Seedance 2.5, Wan 3.0 and MiniMax H3 matter for the numbers they add to this list, not for the headline.
The five questions
| Question | If the answer is | Rows that remain |
|---|---|---|
| How long? | over 15 s | seedance-2.5, wan-3.0 |
| How long? | under 3 s | wan-3.0 |
| How sharp? | 4K | gemini-omni-flash-1.1 |
| Sound? | from the model, always | minimax-h3 rows, gemini-omni-flash-1.1 |
| Input? | a source video to edit | gemini-omni-flash-1.1 (edit), genjutsu, h3-max-recast |
| References? | images, videos and audio | seedance-2.5, wan-3.0, minimax-h3 rows |
Why not rank by quality
Quality depends on your prompt and subject, and the v1 rows take no seed, so a single sample proves little. A capability filter is repeatable. After it, run the same prompt on the surviving models at 480p and compare blind.
Where new models fit
Sume lists Seedance 2.5, Wan 3.0, MiniMax H3 and MiniMax H3 Max, Kling 3 and Gemini Omni Flash 1.1. It does not list an LTX model or a Veo id; Gemini Omni Flash is the Google row. If a model you read about is not in GET /v1/video-router/models, it is not available through the API.
Then pin
Once you know the row, pin the bare id (for example wan-3.0) on POST /v1/videos. Use sume/auto only where you do not care which model ran.
Sources
Related posts
More in Models
- Which AI video models can't do text-to-video? Rows that need a source
Grok Imagine Video 1.5, Genjutsu Motion Transfer and H3 Max Recast refuse a prompt-only request. What each needs, and which rows accept text alone.
- AI video reference limits: how many images and clips per model
Reference image and reference video caps for Wan 3.0, MiniMax H3, Gemini Omni Flash, Genjutsu and H3 Max Recast on Sume, in one table with the odd limits.
- AI voice agent latency budget: who owns which 100 milliseconds
Vendor numbers for speech-to-text, the language model and text-to-speech side by side, with a note on where an async file API like Sume belongs and where not.
- Can you sell images from open-weights models? Licences compared
Open weights do not mean commercial use. Ideogram 4, Qwen-Image, FLUX.2 dev and LTX-2.5 differ on selling outputs. What each page says, and hosted rows.
Written by Sume