AI video reference limits: how many images and clips per model
Reference image and reference video caps for Wan 3.0, MiniMax H3, Gemini Omni Flash, Genjutsu and H3 Max Recast on Sume, in one table with the odd limits.

Reference limits are the most model-specific part of Sume's video API: Wan 3.0 takes up to 10 reference images and 5 reference videos, MiniMax H3 up to 9 images and 3 videos, Gemini Omni Flash 1.1 up to 10 images and 3 videos of 3 seconds each, and the swap rows take 1 to 8 (Genjutsu) or 1 to 4 (Recast) photos. Kling 3 and Grok Imagine take none.
The table
Fields are reference_image_urls, reference_video_urls and reference_audio_urls on the Video Router, and input_references on /v1/videos.
| Model id | Images | Videos | Audio |
|---|---|---|---|
| wan-3.0 | up to 10 | up to 5 (15 s total, 16 fps or higher) | up to 5 (15 s total) |
| minimax-h3, minimax-h3-max | up to 9 | up to 3 (2 to 15 s each, 15 s combined) | up to 3 (2 to 15 s each, 15 s combined) |
| gemini-omni-flash-1.1 | up to 10 | up to 3 (each 3 s or less) | not accepted |
| higgsfield-genjutsu | 1 to 8 | exactly 1 source | not accepted |
| h3-max-recast | 1 to 4 | exactly 1 source | not accepted |
| kling-3, grok-imagine-video-1.5 | none | none | none |
Odd limits that trip requests
These are the failures seen when a request that worked on one model is replayed on another.
- MiniMax H3 counts images, videos and audios together: 12 in total, and audio cannot be the only reference.
- Gemini Omni Flash reference clips are capped at 3 seconds each, much shorter than Wan's 15 seconds total.
- Gemini Omni Flash addresses references in the prompt as
<IMAGE_REF_0>and<VIDEO_REF_0>, zero-based in list order. - A single image with no first or end frame on Gemini Omni Flash is treated as reference-to-video, not image-to-video.
Choosing by reference count
If you have a character sheet of many images, Wan 3.0 and Gemini Omni Flash take 10. If you have motion in a clip you want to borrow, Wan accepts the longest combined reference footage. If you only need to keep one face consistent, one image on any reference-capable row is enough.
Sources
Related posts
More in Models
- AI voice agent latency budget: who owns which 100 milliseconds
Vendor numbers for speech-to-text, the language model and text-to-speech side by side, with a note on where an async file API like Sume belongs and where not.
- Can you sell images from open-weights models? Licences compared
Open weights do not mean commercial use. Ideogram 4, Qwen-Image, FLUX.2 dev and LTX-2.5 differ on selling outputs. What each page says, and hosted rows.
- ChatGPT Try On from a screenshot: the same edit through an API
ChatGPT Try On starts from a selfie plus a product screenshot. Do the same edit with openai/gpt-image-2.5 on Sume: two references, one prompt, one Python call.
- Chinese, Hindi and Russian text to speech API: Sume zh, hi and ru
Sume's Voices library has zh, hi and ru tags. How to request each, what Eleven v4 lists, and the one field that stops an English-sounding read.
Written by Sume