Gemini Omni Flash references on Sume: 10 images, 3 videos of 3 s
Omni Flash 1.1 on Sume takes up to 10 reference images and 3 reference videos of at most 3 s each, so 9 s of video. No audio reference. Tags and costs.

Gemini Omni Flash 1.1 on Sume accepts up to 10 reference images and up to 3 reference videos, each at most 3 s long, so the most video reference one request can carry is 3 x 3 = 9 s. It does not accept audio references. You refer to the media in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>, counting from zero in list order.
The limits in one table
From the Sume Video Router page, read 2026-10-09. The model ID is gemini-omni-flash-1.1, and it is one catalog ID: Sume routes by the shape of your request, and you never choose an endpoint.
| Capability | You send | Limits |
|---|---|---|
| text_to_video | prompt | 3 to 10 s; 360p, 720p, 1080p or 4K; 16:9 or 9:16 |
| image_to_video | image_url, optional end_image_url | Same envelope |
| reference_to_video | reference_image_urls and/or reference_video_urls | Up to 10 images; up to 3 videos, each up to 3 s |
| video_to_video (edit) | video_url | Prompt gives the edit; resolution optional (default 720p); no aspect_ratio or duration |
Cost of a referenced clip
References do not appear as a separate line in the catalog rates: Sume bills per output second by resolution. An 8-second referenced clip is 8 x $0.125 = $1.00 at 720p, $1.50 at 1080p and $3.00 at 4K. Native synced audio is always on, and the API rejects generate_audio: false.
Google's pricing page lists Omni Flash input, including video, at $1.50 per 1M tokens, but the Sume catalog rates this page uses are per output second, so do not add a token estimate on top of the Sume figure.
Prompt tags in practice
Tags follow list order and start at zero. If you send two images and one video, the images are <IMAGE_REF_0> and <IMAGE_REF_1> and the video is <VIDEO_REF_0>. A prompt such as the character in <IMAGE_REF_0> walks through the set shown in <IMAGE_REF_1>, moving like <VIDEO_REF_0> uses all three. Keep the order of the URL lists stable between retries so that the tags keep their meaning.
The 3-second ceiling on each reference video is the practical limit. A 12-second source must be trimmed to three clips of 3 s or fewer. Sume lists a trim endpoint at $0.02 per job for this kind of cut, so trimming a source costs two cents per cut before you spend on generation.
Common mistakes
Do not send video_url together with image_url, end_image_url or any reference_*_urls; video_url is the edit source, not a reference. Do not send reference_audio_urls: the model has no such input, and there is no bitrate_mode. If you need audio references, the Sume docs say the Seedance 2.x models, Wan 3.0, MiniMax H3 and H3 Max accept audio and video references; Omni does not.
Reference media must be reachable over public HTTPS. The errors page lists input_media_unreachable and image_not_fetchable for media Sume could not fetch.
Sources
Related posts
More in Models
- GPT Image 1 retires Oct 23: a 20-prompt test on 2.5 for about $6.06
OpenAI lists gpt-image-1 for removal on October 23, 2026. Test 20 prompts at low, medium and high on gpt-image-2.5 through Sume for about $6.06 first.
- GPT Image 2.5 at 1024px: xhigh is 3,122 output tokens, max is 7,024
Sume's docs give GPT Image 2.5 output estimates at 1024x1024: xhigh $0.09366, max $0.21072 at $30 per million tokens, i.e. 3,122 and 7,024 tokens.
- GPT Image 2.5 quality auto reserves max: set quality yourself on Sume
On Sume, GPT Image 2.5 quality auto reserves max. At 1024px the max output estimate is $0.2634 with margin versus a $0.02475 low row. Set quality explicitly.
- Grok Imagine Video 1.5 on Sume: image-in only, silent, $0.19 for 15 s
Grok Imagine Video 1.5 on Sume needs a start image, makes silent 480p or 720p clips of 4 to 15 seconds, and a 15-second clip is 15 x $0.0125 = $0.1875.
Written by Sume