Gemini Omni reference errors: 10 images, 3 clips of 3 s, 16:9 or 9:16
Gemini Omni Flash 1.1 on Sume refuses an 11th image, a 4th or longer-than-3-second clip, audio references, bitrate_mode and ratios other than 16:9 or 9:16.

gemini-omni-flash-1.1 has the tightest reference rules in the Video Router catalog. It takes at most 10 reference images and 3 reference videos of at most 3 seconds each. It refuses audio references and bitrate_mode, and it accepts only 16:9 or 9:16. Each limit has its own named 400, so the message tells you which one you hit.
Each rule and its message
The messages come from the validation code on main. The pattern is the model name plus the limit, which makes them easy to match in logs.
| You send | Message |
|---|---|
| 11 or more reference_image_urls | model gemini-omni-flash-1.1 accepts at most 10 reference_image_urls. |
| 4 or more reference_video_urls | model gemini-omni-flash-1.1 accepts at most 3 reference_video_urls (each at most 3 seconds). |
| reference_audio_urls | reference_audio_urls is not supported by model gemini-omni-flash-1.1. |
| bitrate_mode | bitrate_mode is not supported by model gemini-omni-flash-1.1. |
| aspect_ratio 4:3, 1:1 and so on | model gemini-omni-flash-1.1 accepts aspect_ratio 16:9 or 9:16. |
| aspect_ratio with video_url (edit) | aspect_ratio is not supported on gemini-omni-flash-1.1 edit (video_url); the output keeps the source framing. |
What the schema cannot see
The count of clips is checked at the API, but the three-second length per clip is a rule stated in the message and the catalog notes, not something the validation can measure on a URL. Measure each reference clip before you upload. If a clip is longer, cut it first: video_trim makes a new MP4 from a start and an end or duration on a stored clip.
Trim, then submit
The snippet shows the trim request body for a 3-second reference cut from a stored clip. It uses the documented video_url plus start and duration fields and prints the body; set SEND=1 to submit it, which creates a billed job. Use your own media.sume.com URL and a fresh idempotency key per request.
import os
import requests
body = {
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"start": 4,
"duration": 3,
}
print(body)
if os.environ.get("SEND") == "1":
r = requests.post("https://api.sume.com/v1/video-trim",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "trim-ref-001"},
json=body, timeout=60)
print(r.status_code)
When another model is better
If you need audio references, longer reference clips or other aspect ratios, use Wan 3.0 or a Seedance model. If you need native 4K, Gemini Omni is one of two models that take it. The Video Router docs table lists the same limits for the model, so the code and docs agree on everything here except the audio flag covered in another post.
One last tip: validate the aspect ratio before you generate any assets for the shot. A square or 4:3 composition made for another model cannot be sent here, and cropping a finished reference to 16:9 or 9:16 afterwards changes what the model sees. Decide the frame shape first, then choose the references to fit it, then check the counts. That order avoids rework when a model has narrow rules like these.
Sources
Related posts
More in Models
- Gemini Omni video in the Windows app, and by API on Sume
Google lists Gemini Omni video in the Gemini app for Windows. For your own app, Sume exposes gemini-omni-flash-1.1 at 360p to 4K, 3 to 10 seconds.
- generate_audio false on Gemini Omni: accepted, clip keeps its audio
Video Router accepts generate_audio false on gemini-omni-flash-1.1 and ignores it; MiniMax H3 answers 400. Check the model before you rely on a mute flag.
- GPT Image 2.5 above 2560x1440 is experimental: a safe size ladder
OpenAI marks sizes over 2560x1440 experimental on GPT Image models. A ladder of valid sizes up to 3840x2160 and a Python check for Sume's image_size.
- GPT Image 2.5 quality: OpenAI defaults to auto, Sume to high
OpenAI's default quality for GPT Image 2.5 is auto; Sume's is high when omitted. Why that differs, what auto reserves on Sume, and how to pin quality.
Written by Sume