AI video reference limits: how many images and clips per model

Reference image and reference video caps for Wan 3.0, MiniMax H3, Gemini Omni Flash, Genjutsu and H3 Max Recast on Sume, in one table with the odd limits.

5 min readSume
All posts

Reference limits are the most model-specific part of Sume's video API: Wan 3.0 takes up to 10 reference images and 5 reference videos, MiniMax H3 up to 9 images and 3 videos, Gemini Omni Flash 1.1 up to 10 images and 3 videos of 3 seconds each, and the swap rows take 1 to 8 (Genjutsu) or 1 to 4 (Recast) photos. Kling 3 and Grok Imagine take none.

The table

Fields are reference_image_urls, reference_video_urls and reference_audio_urls on the Video Router, and input_references on /v1/videos.

Reference input caps, read 2026-10-03. Source: Sume video catalog on origin/main, read 2026-10-03.
Model idImagesVideosAudio
wan-3.0up to 10up to 5 (15 s total, 16 fps or higher)up to 5 (15 s total)
minimax-h3, minimax-h3-maxup to 9up to 3 (2 to 15 s each, 15 s combined)up to 3 (2 to 15 s each, 15 s combined)
gemini-omni-flash-1.1up to 10up to 3 (each 3 s or less)not accepted
higgsfield-genjutsu1 to 8exactly 1 sourcenot accepted
h3-max-recast1 to 4exactly 1 sourcenot accepted
kling-3, grok-imagine-video-1.5nonenonenone

Odd limits that trip requests

These are the failures seen when a request that worked on one model is replayed on another.

  • MiniMax H3 counts images, videos and audios together: 12 in total, and audio cannot be the only reference.
  • Gemini Omni Flash reference clips are capped at 3 seconds each, much shorter than Wan's 15 seconds total.
  • Gemini Omni Flash addresses references in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>, zero-based in list order.
  • A single image with no first or end frame on Gemini Omni Flash is treated as reference-to-video, not image-to-video.

Choosing by reference count

If you have a character sheet of many images, Wan 3.0 and Gemini Omni Flash take 10. If you have motion in a clip you want to borrow, Wan accepts the longest combined reference footage. If you only need to keep one face consistent, one image on any reference-capable row is enough.

Sources

Related posts

More in Models

All Models posts

Written by Sume