AI Video Reference Limits: How Many Images, Clips and Audio Files
How many reference images, videos and audio clips Wan 3.0, Gemini Omni, Veo 3.1 and Seedance 2.5 accept, and where Sume's catalog stops.

Reference-to-video lets you hand a model your product, your character or a motion clip. The catch is that the counts differ sharply, and an over-limit request is rejected, not trimmed. These are the limits as documented on 2026-10-03.
The counts
Sume's limits come from its catalog constraints. Vendor limits come from the vendor's page and may be wider than what Sume exposes.
| Model | Images | Videos | Audio |
|---|---|---|---|
| Wan 3.0 on Sume | up to 10 | up to 5 (15 s total, 16 fps or more) | up to 5 (15 s total) |
| Gemini Omni Flash 1.1 on Sume | up to 10 | up to 3, each 3 s or less | none |
| Kling Video 3 on Sume | none | none | none |
| Seedance 2.5 per ByteDance | up to 30 | up to 10 | up to 10 |
| Veo 3.1 per Google | up to 3 | not listed | not listed |
Reading the table
Sume's catalog marks Seedance rows as supporting image, video and audio references, but its docs do not publish counts for them, so the ByteDance figures are the vendor's and should be treated as an upper bound until you test. Kling 3 on Sume has no reference fields; it accepts a first frame and an end frame only.
Google's Veo page ties reference images to the 8-second duration, which also covers 1080p and 4K. Omni's reference videos are limited to three clips of three seconds each, and audio inside them is ignored.
Using them in a request
On POST /v1/videos, input_references carries reference media and frame_images carries first and last frames. If both are sent, frame_images wins and the request is treated as image-to-video. Omni uses 0-based tokens such as <IMAGE_REF_0> and <VIDEO_REF_0> in the prompt to point at a reference.
Practical rule: send the fewest references that pin the look. Ten product photos from one angle add cost without adding information. See the checklist for pinning a model.
Sources
Related posts
More in Models
- AI video reference limits: how many images and clips per model
Reference image and reference video caps for Wan 3.0, MiniMax H3, Gemini Omni Flash, Genjutsu and H3 Max Recast on Sume, in one table with the odd limits.
- AI Video With Built-In Audio: Which Models Always Render Sound
Which video models always render audio, which have a switch, and which have none: Veo 3.1, Gemini Omni, Kling 3, H3, Grok and Seedance on Sume.
- AI voice agent latency budget: who owns which 100 milliseconds
Vendor numbers for speech-to-text, the language model and text-to-speech side by side, with a note on where an async file API like Sume belongs and where not.
- Arabic text to speech API: send language ar to Sume TTS
Arabic is on Cartesia's Sonic 3.6, 3.5 and 3 but never on Sonic 2. How to send ar through Sume TTS, which voice id to use, and what an Arabic script costs.
Written by Sume