Gemini Omni with several videos: 3 references, no cross-video use
Gemini Omni takes up to 3 reference clips of 3 s each, yet Google warns that reasoning across several videos may degrade output. Sume: one video_url source.

Gemini Omni Flash accepts up to 3 reference clips of up to 3 seconds each, but Google's limitations list also says that referencing or reasoning across multiple videos is not supported and may give degraded or unexpected output. Read together: give each clip its own job in the prompt, such as one character per clip, rather than asking the model to compare, merge or choose between videos.
Both lines are from Google's Gemini API: Generate and edit videos with Gemini Omni Flash, read 2026-10-02. Sume's Video Router mirrors the clip limit: reference_video_urls takes at most 3 clips of at most 3 seconds for gemini-omni-flash-1.1, and the edit source video_url is a single clip.
Are those two Google lines in conflict?
They are different things on the page. The 3 clips by 3 seconds line is an input limit on video references; the second is a warning about multi-video prompting. The page does not define the boundary between them, so treat a single clear role per reference as the safe use and test anything beyond it.
How do I use more than one reference safely?
Google's tag guide gives an example that binds one video as a character reference and one image as an object reference: The woman in <VIDEO_REF_0> is playing the violin shown in <IMAGE_REF_0>. That keeps a single video in play. On Sume the same addressing works inside the prompt, with <VIDEO_REF_0> for the first entry in reference_video_urls, counted from 0 in list order.
The example below follows the Video Router page's reference request, with one clip and one image.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-ref-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "A cat inspired by <IMAGE_REF_0> walks through the setting shown in <VIDEO_REF_0>.",
"reference_image_urls": ["https://example.com/inputs/cat.png"],
"reference_video_urls": ["https://example.com/inputs/setting-3s.mp4"],
"resolution": "720p",
"duration": 8,
"aspect_ratio": "16:9",
"mode": "async"
}'What are the limits side by side?
Sume's schema allows more reference clips on other models, but names 3 for Omni.
| Item | Sume | |
|---|---|---|
| Reference clips | Up to 3 clips, 3 seconds each | reference_video_urls: at most 3, each at most 3 seconds |
| Reasoning across several videos | Not supported; may degrade output | Not described in the docs |
| Edit source | One video, 10 seconds or less when uploaded | video_url, one clip |
| Reference audio | Ignored in video references | Omni takes no reference_audio_urls |
What would I check first if a multi-clip prompt looks wrong?
Start by simplifying the input.
- Cut to one reference clip and see if the result improves.
- Check each clip is under 3 seconds.
- Make sure the prompt names every
<VIDEO_REF_N>it expects the model to use.
Sources
Related posts
More in Models
- Gemini Omni REST: output_video is SDK-only, read the steps array
Calling Gemini Omni over REST? interaction.output_video is SDK-only. Read the base64 video from the model_output step; on Sume you get a media URL instead.
- Gemini Omni [# Sources] and [# References] tags vs Sume fields
Gemini Omni binds media to roles with tags like <FIRST_FRAME> and [# Sources ...]. Sume uses request fields instead: image_url, end_image_url, reference lists.
- Gemini Omni and Veo in Korean: only English is fully supported
Google's Omni and Veo pages say English is fully supported; other languages aren't evaluated. For Korean prompts, describe in English and quote on-screen text.
- Gemini Omni can't use a YouTube link as its source video
Google lists YouTube videos as an unsupported media source for Gemini Omni Flash. On Sume, video_url is a URI field; use a direct link to the clip file.
Written by Sume