Gemini app video: one video and 5 images in, vs Sume's 10 references
The Gemini app takes one video and up to 5 images per generation. Sume's Gemini Omni Flash 1.1 takes up to 10 image references and 3 short video references.

In the Gemini app you can upload one video and up to 5 images for a single generation. On Sume, the Gemini Omni Flash 1.1 reference mode accepts up to 10 image URLs and up to 3 video URLs of 3 seconds or less each, but the docs list a video edit as the source clip plus a prompt, with no references.
Both sides are read from first-party pages on 2026-10-02: the Gemini Apps Help page and the Gemini API Omni guide for Google, and the Sume Video Router docs. The numbers are not the same thing, so a project that fits one will not always fit the other.
What does the Gemini app let me upload?
Google's help page lists these requirements and limits for video in the app: a Google AI plan for personal accounts, a qualifying Workspace license for work or school accounts, an age of 18 or over, and being signed in to Gemini Apps. You can upload one video and up to 5 images per generation. Text prompts default to landscape and you can change the aspect ratio.
Multi-turn editing works inside one conversation. The page also says video takes a few minutes and that you cannot interact with the same chat while a video generates, though you can start a new chat. Video uploads for edits are unavailable in the EEA, Switzerland, the United Kingdom and some US states.
What does the Gemini API take?
The API guide is the stricter of the two for video input. An input video is at most 10 seconds unless you are extending through a multi-turn chain, and video references are at most three clips of three seconds each. The guide's task table lists text to video, image to video, first and last frame, subject reference with two or more images, edit and extend. It does not give a hard image count for subject reference, so do not quote one for the API.
How does Sume map to those?
Two details trip people up. In the app, one video plus images can be sent together. On Sume, an edit and a reference run are different requests.
- Address references in the prompt as
<IMAGE_REF_0>and<VIDEO_REF_0>, counted from zero in list order. - Sume has no
reference_audio_urlsfor this model. video_urlis the edit source, not a reference, and the docs list no reference inputs for an edit.
| What you want | Gemini app | Sume Video Router |
|---|---|---|
| Text to video | Prompt, landscape by default | prompt, 16:9 or 9:16, 3 to 10 seconds |
| Reference images | Up to 5 per generation | reference_image_urls, up to 10 |
| Reference video | One uploaded video | reference_video_urls, up to 3, each 3 seconds or less |
| Edit a video | Upload one video and instruct | video_url plus a prompt; no references listed |
| Audio | Generated with the video | Always on; generate_audio: false is rejected |
Which should I use for a 7-image product shot?
If you have seven product photos, the app's 5-image limit forces you to choose. Sume's router accepts all seven as reference_image_urls. Name each one in the prompt so the model knows which item goes where, and keep the clip between 3 and 10 seconds.
Say honestly what that buys you: a higher count of accepted inputs is not a guarantee that the model will use every one. Run a 360p draft, check that each product appears, and only then render at 1080p.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: seven-shots-draft-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "A tabletop shot. <IMAGE_REF_0> slides in, then <IMAGE_REF_1>. Calm music.",
"reference_image_urls": ["https://example.com/a.png", "https://example.com/b.png"],
"duration": 6,
"resolution": "360p",
"aspect_ratio": "9:16",
"mode": "async"
}'How do I plan a clip around these limits?
Work backwards from the shot list. If a scene needs more than five distinct inputs, split it into two clips and join them later; Sume's Timeline joins ordered video slots into one MP4, and the Gemini app can share to YouTube or download the result. If a scene needs a short moving reference, such as a three-second gesture, remember that both Google's API and Sume cap each video reference at three seconds and Sume allows three of them.
Keep each reference doing one job. One image for the character, one for the product, one for the setting reads better in a prompt than five near-duplicates of the same item. On Sume, name them in order, <IMAGE_REF_0> first, and say what each one is in plain words so the sentence still makes sense to a person reading the request later.
Finally, budget for retries. A draft at 360p costs a fraction of a full render, according to Google's launch post, which says drafts run up to 60 percent faster at a third of the cost of standard 720p. Use that to find the prompt that uses your references the way you meant, then spend on the final size.
Sources
Related posts
More in Comparisons
- Gemini batch create is not idempotent: two jobs vs Sume bulk replay
Gemini's docs say sending the same batch creation request twice creates two batch jobs. Sume bulk runs replay the old queue for the same Idempotency-Key.
- Where is Gemini Omni 1.1 Flash? Google's named partners and Sume
Google names Adobe Firefly, Figma Weave, Runway and GMI Cloud as production users of Omni 1.1 Flash. What that means if you want one API with job ids.
- Gemini Omni Flash for client work: app, Flow, Shorts, or an API?
Gemini Omni Flash is in five places. Who each is for, what is free, and when a client project needs an API with job ids rather than an app.
- Gemini Omni won't take photos of minors in the EEA, UK or Switzerland
Google's Omni API guide blocks uploading images with minors in the EEA, UK and Switzerland. What the guide says, what Sume's docs say, and what to check first.
Written by Sume