Reference limits per Sume video model: images, clips, audio

How many reference images, videos and audio clips each Sume video model takes on /v1/videos: Seedance 12 total, Wan 10/5/5, H3 9/3/3, Omni 10 and 3.

5 min readSume
All posts

Sume's /v1/videos handler caps references per model: Seedance ids take 12 in total, Wan 3.0 takes 10 images, 5 videos and 5 audio clips, MiniMax H3 and H3 Max take 9, 3 and 3 with a total of 12, and Gemini Omni Flash 1.1 takes 10 images and 3 clips of at most 3 seconds each. Kling 3.0 and Grok Imagine 1.5 take no references.

The numbers below are read from the handler and from the Video Router docs on 2026-10-02. Use them to pick a model before you build a prompt around nine references that the model will refuse.

What does each model accept?

Counts are for input_references entries on /v1/videos. Read the same capability flags for the Video Router from GET /v1/video-router/models.

Reference caps by Sume video model, read 2026-10-02
Model idImagesVideosAudioOther rule
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini12 total across all types(in the 12)(in the 12)Sume checks the 12 total; fal's Seedance 2.0 page also lists 9 images, 3 videos and 3 audio files. Clip 4-30 s on 2.5, 4-15 s on the 2.0 ids
wan-3.01055Catalog: videos 15 s total and 16 fps or more; audio 15 s total
minimax-h3, minimax-h3-max93312 total
gemini-omni-flash-1.1103, each 3 s or lessnot acceptedNative audio always on
higgsfield-genjutsu1 to 8exactly 1 sourcenot acceptedMotion transfer only
h3-max-recast1 to 4exactly 1 sourcenot acceptedOne photo per person
kling-3, grok-imagine-video-1.5nonenonenoneUse frame images

Why do the numbers differ so much?

Each model's own provider sets the envelope, and Sume checks it before submitting. Sume caps Seedance at 12 in total; Wan's limits are split by type; Gemini Omni limits the length of each reference clip to 3 seconds. The catalog is the contract: capabilities on GET /v1/video-router/models and supported_input_references on GET /v1/videos/models list which types a model accepts, and the error messages state the count.

Two of the rows are not general reference models at all. Genjutsu and Recast take a source video plus photos and rewrite the source, and they are explicit picks that sume/auto never routes to.

A rule of thumb from the table: Wan 3.0 and Gemini Omni take the most images (10 each), Seedance takes 12 files in total, and if you need more than 3 audio clips, Wan 3.0 is the one with a Sume-stated audio allowance above 3. Anything beyond those numbers is a split into several clips, joined afterwards.

Which model should you pick for many references?

For a prompt that uses one product shot, one character and one scene, any of the Seedance ids works and the 12 total is generous. For a cast of several characters plus a style clip and a voice clip, Wan 3.0's 10, 5 and 5 give the most headroom by type, and it is the only row where Sume states a per-type audio cap above 3.

If the references are short motion clips, remember the sub-limits: Wan's 15 seconds combined, Gemini Omni's 3 seconds each. A long reference video fits none of those; for Seedance, check the provider's own clip-length limits (fal's Seedance 2.0 page lists 2 to 15 seconds combined for reference video).

How do you read the limits at runtime?

Call the catalog and read the model row rather than hard-coding the table, because limits move. GET /v1/video-router/models/{model_id} returns one row and a 404 for an unknown id; the capability flags reference_images, reference_videos and reference_audios tell you which types exist. The numeric caps are in the error text, so a validation request that is expected to fail is the cheapest probe.

For the Seedance-specific arithmetic see the 12-reference post, and for which models take a video reference at all see the reference video post.

A practical pattern is to keep one probe request per model in your test suite: send one reference more than the cap and assert on unsupported_capability. When the error text changes, you learn the limit moved before a customer does. Because validation happens before the provider call, the probe is a cheap request and needs no real media beyond valid HTTPS URLs.

Sources

Related posts

More in Models

All Models posts

Written by Sume