Reference limits per Sume video model: images, clips, audio
How many reference images, videos and audio clips each Sume video model takes on /v1/videos: Seedance 12 total, Wan 10/5/5, H3 9/3/3, Omni 10 and 3.

Sume's /v1/videos handler caps references per model: Seedance ids take 12 in total, Wan 3.0 takes 10 images, 5 videos and 5 audio clips, MiniMax H3 and H3 Max take 9, 3 and 3 with a total of 12, and Gemini Omni Flash 1.1 takes 10 images and 3 clips of at most 3 seconds each. Kling 3.0 and Grok Imagine 1.5 take no references.
The numbers below are read from the handler and from the Video Router docs on 2026-10-02. Use them to pick a model before you build a prompt around nine references that the model will refuse.
What does each model accept?
Counts are for input_references entries on /v1/videos. Read the same capability flags for the Video Router from GET /v1/video-router/models.
| Model id | Images | Videos | Audio | Other rule |
|---|---|---|---|---|
| seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini | 12 total across all types | (in the 12) | (in the 12) | Sume checks the 12 total; fal's Seedance 2.0 page also lists 9 images, 3 videos and 3 audio files. Clip 4-30 s on 2.5, 4-15 s on the 2.0 ids |
| wan-3.0 | 10 | 5 | 5 | Catalog: videos 15 s total and 16 fps or more; audio 15 s total |
| minimax-h3, minimax-h3-max | 9 | 3 | 3 | 12 total |
| gemini-omni-flash-1.1 | 10 | 3, each 3 s or less | not accepted | Native audio always on |
| higgsfield-genjutsu | 1 to 8 | exactly 1 source | not accepted | Motion transfer only |
| h3-max-recast | 1 to 4 | exactly 1 source | not accepted | One photo per person |
| kling-3, grok-imagine-video-1.5 | none | none | none | Use frame images |
Why do the numbers differ so much?
Each model's own provider sets the envelope, and Sume checks it before submitting. Sume caps Seedance at 12 in total; Wan's limits are split by type; Gemini Omni limits the length of each reference clip to 3 seconds. The catalog is the contract: capabilities on GET /v1/video-router/models and supported_input_references on GET /v1/videos/models list which types a model accepts, and the error messages state the count.
Two of the rows are not general reference models at all. Genjutsu and Recast take a source video plus photos and rewrite the source, and they are explicit picks that sume/auto never routes to.
A rule of thumb from the table: Wan 3.0 and Gemini Omni take the most images (10 each), Seedance takes 12 files in total, and if you need more than 3 audio clips, Wan 3.0 is the one with a Sume-stated audio allowance above 3. Anything beyond those numbers is a split into several clips, joined afterwards.
Which model should you pick for many references?
For a prompt that uses one product shot, one character and one scene, any of the Seedance ids works and the 12 total is generous. For a cast of several characters plus a style clip and a voice clip, Wan 3.0's 10, 5 and 5 give the most headroom by type, and it is the only row where Sume states a per-type audio cap above 3.
If the references are short motion clips, remember the sub-limits: Wan's 15 seconds combined, Gemini Omni's 3 seconds each. A long reference video fits none of those; for Seedance, check the provider's own clip-length limits (fal's Seedance 2.0 page lists 2 to 15 seconds combined for reference video).
How do you read the limits at runtime?
Call the catalog and read the model row rather than hard-coding the table, because limits move. GET /v1/video-router/models/{model_id} returns one row and a 404 for an unknown id; the capability flags reference_images, reference_videos and reference_audios tell you which types exist. The numeric caps are in the error text, so a validation request that is expected to fail is the cheapest probe.
For the Seedance-specific arithmetic see the 12-reference post, and for which models take a video reference at all see the reference video post.
A practical pattern is to keep one probe request per model in your test suite: send one reference more than the cap and assert on unsupported_capability. When the error text changes, you learn the limit moved before a customer does. Because validation happens before the provider call, the probe is a cheap request and needs no real media beyond valid HTTPS URLs.
Sources
Related posts
More in Models
- Swap the product in a UGC clip with Gemini Omni Flash 1.1 video edit
Keep a UGC clip that works and change only the product: Sume's video router edit mode for gemini-omni-flash-1.1 takes a video_url and a one-line prompt.
- Veo 3.1 and Veo 3.1 Fast previews end October 22: what replaces them
Google's deprecations page sets October 22, 2026 as the shutdown date for the Veo 3.1 and 3.1 Fast previews. Dates, the replacement id and what Sume lists.
- Veo seed doesn't make output repeatable; Sume rejects seed
Google says Veo's seed only slightly improves determinism. Sume's v1 video models report seed false and reject the field; keep the output file to get a repeat.
- Which AI video model gives 1080p on Sume, and which stop at 768p?
Sume lists 1080p for Seedance, Wan 3.0, Kling 3.0 and Auto; MiniMax H3 is native 480p or 768p in the panel. Resolution table by model, with the API check.
Written by Sume