Omni video references: three clips of 3 s, so 9 s of footage at most
Gemini Omni Flash 1.1 takes up to 3 reference videos of 3 s each and 10 reference images. The tags are VIDEO_REF_0 to 2; the price is unchanged.

Gemini Omni Flash 1.1 on Sume accepts up to three reference videos, each no longer than three seconds, so you can hand it at most nine seconds of reference footage per request. You refer to them in the prompt as <VIDEO_REF_0>, <VIDEO_REF_1> and <VIDEO_REF_2>, numbered from zero in the order of the list.
The limits
These limits come from the Video Router docs and the catalog constraints for the model. Reference media is sent before the prompt, and the tag numbers follow the list order.
| Field | Maximum | Tag in the prompt | Note |
|---|---|---|---|
| reference_image_urls | 10 | <IMAGE_REF_0> to <IMAGE_REF_9> | 0-based, list order |
| reference_video_urls | 3 | <VIDEO_REF_0> to <VIDEO_REF_2> | each clip at most 3 s |
| reference_audio_urls | not accepted | none | the model takes no audio input |
A request with two clips
The body below sends two reference clips and one image to POST /v1/video-router/generate. The prompt names each by its tag, so the model knows which asset you mean.
Trim long source footage to 3 seconds or less before you send it. If your clip is longer, cut a 3 s piece with the Video trim tool first; the Omni request does not accept a longer reference.
{
"model": "gemini-omni-flash-1.1",
"prompt": "Keep the motion of <VIDEO_REF_0>, and the mug from <IMAGE_REF_0>.",
"reference_image_urls": ["https://example.com/mug.png"],
"reference_video_urls": [
"https://example.com/pan.mp4",
"https://example.com/steam.mp4"
],
"resolution": "720p",
"duration": 6,
"aspect_ratio": "9:16",
"mode": "async"
}Cost with references
Billing follows the output length and the resolution, not the number of references. The request above is a 6 s clip at 720p: 6 x 12.5 = 75 cents. The same 6 s clip without any reference is also 75 cents.
Native synced audio is always on, and Omni does not take reference audio, so a voice or a track has to be added after the render, for example in a Timeline.
The same length at every tier
For reference, a 6-second Omni Flash clip at each resolution. Every price is the seconds times the billable rate, rounded up to a whole cent, as of 2026-10-08.
| Resolution | Arithmetic | Billed |
|---|---|---|
| 360p | 6 x 3.75 = 22.5 cents | $0.23 |
| 720p | 6 x 12.5 = 75 cents | $0.75 |
| 1080p | 6 x 18.75 = 112.5 cents | $1.13 |
| 4K | 6 x 37.5 = 225 cents | $2.25 |
Limits to remember
These apply to every request on this page, from the Video Router and Video generation docs:
- Length is 3 to 10 whole seconds in generation modes; an edit takes no duration.
- Aspect ratio is 16:9 or 9:16; an edit takes no aspect ratio.
- Native synced audio is always on, and
generate_audio: falseis rejected. - There is no
bitrate_mode, no reference audio and noseed. - Billing is the provider list times 1.25 per output second, reserved at submit and shown in
usage.cost.
Sources
Related posts
More in Developers
- One idempotency key per model when a video fallback changes payload
Reusing an Idempotency-Key after switching video models returns 409 idempotency_conflict. Build the key from order, model, and prompt hash so retries stay safe.
- Replace videos.create_and_poll with a requests helper on Sume
The OpenAI Python SDK's create_and_poll and download_content have no Sume twin. Here is a 25-line requests helper with the same call shape and a safe retry key.
- OpenRouter video client on Sume: swap the base URL, change webhooks
A client written for OpenRouter /videos works on Sume after you change the base URL, key and model ids. The webhook body and signature are different.
- Can I use my own LoRA on a hosted video API? Sume has no LoRA field
Sume's /v1/videos has no LoRA, seed or provider-option fields. MiniMax's H3 license allows LoRAs on the open weights. What that means for a character or style.
Written by Sume