Omni Flash references: 10 images, 3 clips of 3 s, 9 s of footage
Gemini Omni Flash 1.1 on Sume takes up to 10 reference images and 3 reference clips of 3 s each (9 s of footage), with no audio reference. Cost: $1.00 for 8 s.

Gemini Omni Flash 1.1 takes at most 10 reference images (reference_image_urls) and at most 3 reference videos (reference_video_urls), each no longer than 3 seconds. The most reference footage a job can carry is therefore 9 seconds. It takes no audio reference, and native audio is always on in the output.
Prices on this page are Sume list prices as of 2026-10-08: the provider list price times 1.25, computed from Sume's pricing package and billed per output second, linear in duration inside each model's valid range. The model ids, ranges and input types come from the Video generation docs and the Video Router docs; the per-second numbers can be cross-checked in the public catalog.
Limits next to price
The references do not change the per-second rate: the docs say Sume bills the provider list × 1.25 per output second as a function of resolution. An 8-second reference job at 720p is 8 s × $0.125 = $1.00, the same as a text-only 8 s job.
| Item | Limit or price |
|---|---|
| Reference images | up to 10, named <IMAGE_REF_0> to <IMAGE_REF_9> |
| Reference videos | up to 3, each at most 3 s, named <VIDEO_REF_0> to <VIDEO_REF_2> |
| Reference audio | not accepted |
| Output duration | 3 to 10 s |
| Aspect ratios | 16:9 or 9:16 |
| 8 s at 720p | $1.00 |
| 10 s at 1080p | $1.875 |
| 10 s at 4K | $3.75 |
Writing the prompt
Refer to each file in the prompt by its placeholder, zero-based and in list order: <IMAGE_REF_0> is the first image URL you listed, <VIDEO_REF_1> is the second clip. A prompt that says only 'use the references' gives the model no way to tell them apart.
Compare Seedance 2.x, Wan 3.0, MiniMax H3 and H3 Max, which the docs say accept audio and video references; if you need the output to follow a sound track, pick one of those, not Omni.
Scaling this plan up or down
As a yardstick, the plan above is built on Gemini Omni Flash 1.1 at 720p, $0.125 per second. Each extra 5 seconds adds $0.625, a further $10 of budget buys 80 more whole seconds, and the largest single job the model accepts (10 s) holds $1.25 at submit. The shortest one (3 s) holds $0.375.
Those three numbers are enough to rescale the plan without a new table. If the plan doubles, double the totals; if a clip is shortened, subtract the seconds multiplied by the rate; and if the tier changes, swap the rate for the one in the tables above.
- Per second: $0.125
- Per 5 s: $0.625
- Per 10 s: $1.25
- Per 30 s or the model maximum (10 s): $1.25
Check the model's fields first
After the first job completes, read usage.cost on the poll response and compare it with the figure in this post. They should agree, because both are the provider list price times 1.25. If your number differs, the usual cause is a different tier or duration than you planned, not a price change; the public catalog at https://api.sume.com/v1/catalog shows the current rates and needs no key.
Submitting
Send the fields through POST /v1/video-router/generate (flat reference_image_urls and reference_video_urls) or, on /v1/videos, as input_references. Include an Idempotency-Key on the Video Router route so a retry returns the original job.
Sources
Related posts
More in Developers
- Omni resolution values: 4K, lowercase 4k, and what is rejected
Gemini Omni Flash 1.1 on Sume takes 360p, 720p, 1080p and 4K. The catalog also accepts lowercase 4k as an alias. Which strings fail, and the price of each.
- Omni video references: three clips of 3 s, so 9 s of footage at most
Gemini Omni Flash 1.1 takes up to 3 reference videos of 3 s each and 10 reference images. The tags are VIDEO_REF_0 to 2; the price is unchanged.
- One idempotency key per model when a video fallback changes payload
Reusing an Idempotency-Key after switching video models returns 409 idempotency_conflict. Build the key from order, model, and prompt hash so retries stay safe.
- Replace videos.create_and_poll with a requests helper on Sume
The OpenAI Python SDK's create_and_poll and download_content have no Sume twin. Here is a 25-line requests helper with the same call shape and a safe retry key.
Written by Sume