Which AI video model should I use? A checklist by input and output
Pick an AI video model by what you must send and get back: frames, references, audio, length, resolution. This checklist maps each need to Sume catalog ids.

Choose an AI video model by the constraint you cannot bend, not by a leaderboard: whether you have a first or last image, whether you need reference images, video or audio, how long the clip must be, the resolution and whether sound is built in. Each row below names the Sume catalog ids that list that need. This is a how-to-choose list, not a ranking; run one short clip on two candidates before committing.
Everything here is from the Video generation docs and catalog code, read 2026-09-29. Ask GET /v1/videos/models for the live list.
Which model fits which need?
| If you need | Look at |
|---|---|
| A clip longer than 15 s | seedance-2.5 (4 to 30 s), wan-3.0 (2 to 30 s) |
| A start image only | grok-imagine-video-1.5 (image-to-video, 4 to 15 s, no end frame) |
| First and last frame | kling-3, seedance-2.5, seedance-2, wan-3.0, minimax-h3, gemini-omni-flash-1.1 |
| Reference video or audio | Seedance 2.x, wan-3.0, minimax-h3, minimax-h3-max |
| 4K | gemini-omni-flash-1.1, or sume/auto |
| Sound made with the picture | Most catalog models; grok-imagine-video-1.5 has none |
| 4:3 or 3:4 shapes | wan-3.0 lists 16:9, 4:3, 1:1, 3:4, 9:16 |
| No decision | sume/auto |
What should I check first?
- Length: durations are whole seconds and outside the list the request is refused.
- Inputs:
input_referencesare honored only by models that list them;kling-3lists none. - Resolution: list values differ per model (
grok-imagine-video-1.5stops at 720p). - Audio: some models always add sound,
grok-imagine-video-1.5adds none, others followgenerate_audio.
How do I read the live list?
One call returns each model's supported_durations, supported_aspect_ratios, supported_input_references and generate_audio.
curl "https://api.sume.com/v1/videos/models" \
-H "Authorization: Bearer $SUME_API_KEY" \
| jq '.data[] | {id, supported_durations, supported_input_references, generate_audio}'What about price?
Sume reserves provider list × 1.25 on submit, plus a 5.5% default agent fee, and the amounts differ a lot per model. Price a clip with your real length and resolution, not a per-second headline.
Sources
Related posts
More in Models
- Voice cloning API: clone once in the app, speak via the API
Sume has no API route that creates a voice clone. You clone once in the app, then send the voi_ id to the text to speech API on every call.
- Wan 3.0 1080p: how to request full HD from the API
Wan 3.0 outputs 480p, 720p or 1080p. Send resolution 1080p to wan-3.0 on POST /v1/videos, up to 30 seconds, with audio. See the request and the price.
- Wan 3.0 2-second clips: the shortest video length and what it costs
Wan 3.0's floor is 2 seconds, shorter than the 4-second minimum on Seedance 2.x models. Send duration 2 to wan-3.0; at 720p it costs $0.25 before the agent fee.
- Wan 3.0 30-second video: one clip, no stitching
Wan 3.0 makes a single clip of up to 30 seconds with audio. Set duration to 30 on wan-3.0 in POST /v1/videos. Here is the request, the cost, and the limits.
Written by Sume