Gemini Omni Flash 1.1 prompt length on Sume: a 20,000-character check
The Sume catalog constraint for Gemini Omni Flash 1.1 is a prompt of at most 20,000 characters, 3 to 10 seconds. A short Python preflight check before you pay.

gemini-omni-flash-1.1 on Sume takes prompts of up to 20,000 characters, counted as characters and not tokens. Pair that with a duration of 3 to 10 seconds and an aspect ratio of 16:9 or 9:16, and you have the three fields that most often cause a refused request on this row. A local check before you submit costs nothing.
The constraints in one table
These come from the Video Router docs and the constraints returned for the model. Reservation uses the rate in the last row for the duration you send.
| Field | Limit |
|---|---|
| prompt | at most 20,000 characters |
| duration | 3 to 10 seconds, whole seconds |
| aspect_ratio | 16:9 or 9:16 |
| resolution | 360p, 720p, 1080p, 4K |
| generate_audio | not allowed as false; audio is always on |
| reference_image_urls | up to 10 |
| reference_video_urls | up to 3, each at most 3 s |
| rate (10 s, 720p) | 10 x $0.125 = $1.25 |
A preflight function
The function returns a list of problems. Run it before the POST so a long storyboard prompt does not fail after you have waited in a queue.
def omni_problems(prompt, duration, aspect_ratio, generate_audio=None):
out = []
if len(prompt) > 20000:
out.append(f"prompt is {len(prompt)} characters; the limit is 20000")
if not 3 <= duration <= 10:
out.append("duration must be 3 to 10 seconds")
if aspect_ratio not in ("16:9", "9:16"):
out.append("aspect_ratio must be 16:9 or 9:16")
if generate_audio is False:
out.append("audio is always on; generate_audio false is rejected")
return out
print(omni_problems("x" * 20001, 12, "1:1", False))Using the length well
Twenty thousand characters is far more than a shot needs. Long prompts help when you describe several timed beats in one 10-second clip, or when you need to name each reference with the <IMAGE_REF_0> and <VIDEO_REF_0> tokens. For a single shot, a few hundred characters of subject, motion, camera and light are usually enough.
In edit mode the prompt carries the edit instruction, for example replacing one object and keeping the rest unchanged. The same 20,000-character limit applies.
Where the check belongs
Put the function at the edge of your service, right before the POST. A rejected request still costs a round trip and a user's patience, while a local check costs microseconds. Return the list of problems to the caller so a human can fix the prompt.
Count characters the way your language does. In Python len counts code points, which is what you want for a prompt of ordinary text. If your pipeline adds a system preamble or inlines reference tokens, count after assembly.
Sources
Related posts
More in Developers
- Omni Flash defaults to 720p and 16:9: set both on every request
Google's Omni docs default to 720p and 16:9; Sume's Auto defaults to 720p and 8 seconds. Why to send resolution, ratio and duration every time, with prices.
- Gemini Omni edit mode on Sume: video_url only, 720p, no aspect_ratio
Omni edit mode takes a prompt and one video_url. Resolution defaults to 720p, aspect_ratio is rejected, and no image or reference field may be added.
- Omni reference-to-video with 3 clips of 3 s: a 10 s output is $1.25
Gemini Omni Flash 1.1 takes up to 10 reference images and 3 reference clips of 3 seconds each. A 10-second 720p output is 10 x $0.125 = $1.25 on Sume.
- generation_limits missing from a Sume submit: send one, then re-read
The docs say the snapshot is included when Sume can compute it. A Python rule for waves when it is absent, full, or open, tested on three inputs.
Written by Sume