Omni Flash took 7 image references; Omni 1.1 on Sume takes up to 10

Google's original Omni Flash allowed up to 7 image references and 3 short clips. Sume's Omni Flash 1.1 route lists up to 10 images and 3 clips of 3 s.

4 min readSume
All posts

Google's original Gemini Omni Flash listed 3 to 10 second clips at 720p, up to 7 image references and up to 3 clips of 3 seconds or less (Google Cloud, read 2026-10-05). Omni 1.1 Flash added video references up to 3 seconds and scene extension (Google, read 2026-10-05). The Sume Video Router docs for gemini-omni-flash-1.1 list up to 10 image references and up to 3 video references of 3 seconds each. If you found the "7 images" number and wonder which applies on Sume, the Sume number for this id is 10; the 7 belongs to Google's original-model description.

The numbers in one table

Do not carry one source's count to another route. Vendor numbers describe the vendor's own product, while a hosted route documents what it validates.

Omni reference limits by source, read 2026-10-05
SourceModelImage refsVideo refsClip length
Google Cloud blogOriginal Omni FlashUp to 7Up to 3 clips of 3 s or less3 to 10 s, 720p
Google blogOmni 1.1 FlashNot covered hereUp to 3 sExtension to 40 s in 10 s steps
Sume docsgemini-omni-flash-1.1Up to 10 (reference_image_urls)Up to 3, each 3 s or less (reference_video_urls)3 to 10 s

How references are addressed

Sume documents positional tags for Omni 1.1: refer to the media as <IMAGE_REF_0> and <VIDEO_REF_0>, 0-based, in list order. That makes the prompt depend on list order, so keep the order stable.

  • Put the subject image first and call it <IMAGE_REF_0> in the prompt.
  • Use video references for motion or camera style, each no longer than 3 seconds.
  • There is no reference_audio_urls on this id, and native audio is always on.
curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: omni-refs-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "<IMAGE_REF_0> walks toward the camera, matching the camera move in <VIDEO_REF_0>",
    "reference_image_urls": ["https://example.com/subject.png"],
    "reference_video_urls": ["https://example.com/move-3s.mp4"],
    "resolution": "720p",
    "duration": 6,
    "aspect_ratio": "16:9",
    "mode": "async"
  }'

Using ten references well

More references do not mean better output. Ten images of the same product from slightly different angles give the model enough to hold the shape; ten unrelated images give it a mess to reconcile. Group your images by role: one to three for the subject, one for the setting and one for the style. Keep the rest out.

When something goes wrong, remove references before you rewrite the prompt. Cutting from ten to four often shows which image was causing the problem, and it saves cost.

  • Subject first, so it is <IMAGE_REF_0>.
  • Group by role: subject, setting, style.
  • Remove references one at a time when debugging.

A reference list that works

For a product clip, a list of six works well: three product angles, one hand holding it, one table surface and one lighting reference. Name them in the prompt by role, with the product as <IMAGE_REF_0>. If the result ignores the hand, move it earlier in the list and mention it by its tag. Keep the prompt short; the references carry the detail, and the prompt carries the action and the camera.

Remember that references guide the model and do not fix exact frames. If a first frame must match a picture exactly, use image_url instead, which Sume documents for Omni image-to-video.

Read capabilities from GET /v1/video-router/models for the live envelope. Sources: Google Cloud, Google blog, Sume Video Router docs.

Sources

Related posts

More in Models

All Models posts

Written by Sume