Pin Gemini Omni Flash 1.1 or send sume/auto: six checks
sume/auto picks the model for video and defaults to 720p and 8 s. Pin gemini-omni-flash-1.1 when you need edit mode, 4K or fixed references. A decision table.

Send sume/auto when you only need a clip of 3 to 10 seconds at 16:9 or 9:16 and do not care which model makes it; pin gemini-omni-flash-1.1 when the request depends on something specific to Omni: edit mode through video_url, the 4K tier, or the <IMAGE_REF_0> tags. The docs name Omni Flash as the Auto default for video, and Auto answers with sume/auto and never tells you which family ran.
Six checks
Go through the rows from the top. The first one that applies decides.
| Need | Send | Why |
|---|---|---|
| Edit an existing clip | Pin gemini-omni-flash-1.1 on /v1/video-router/generate | The docs describe edit as the video_url field of Video Router |
| 4K output | Pin Omni | The Omni tier list is 360p to 4K; the Auto docs describe a 720p default |
| Tags like <IMAGE_REF_0> | Pin Omni | The tag syntax is documented for this model |
| Same model in an A/B test | Pin | Auto responses show sume/auto and hide the family |
| Longer than 10 s, or 1:1 | Pin another model | Omni is 3 to 10 s, 16:9 or 9:16 |
| Any clip of 3-10 s at 720p | sume/auto | Auto controls default to 720p and 8 s |
What Auto gives you
Auto resolves from the normalized request and the catalog version, so an idempotent replay gets the same price and the same route. The poll response says sume/auto. Sume does not disclose the family and the docs tell you not to guess it from the output.
That is useful for a pipeline that wants no model names. It is a poor fit for reporting that must say which model produced which clip.
{
"model": "sume/auto",
"prompt": "A vertical UGC-style product clip on a desk, natural light",
"aspect_ratio": "9:16",
"duration": 5
}Cost does not decide it
On Omni, 5 s at 720p is 63 cents (62.5 rounded up) and 8 s is $1.00. The same prices apply whether you pin or use Auto when Auto uses Omni, so the choice is about control and reporting, not about saving money. If you need the cheapest possible draft, set resolution to 360p, which is 3.75 cents a second.
The same length at every tier
For reference, a 8-second Omni Flash clip at each resolution. Every price is the seconds times the billable rate, rounded up to a whole cent, as of 2026-10-08.
| Resolution | Arithmetic | Billed |
|---|---|---|
| 360p | 8 x 3.75 = 30 cents | $0.30 |
| 720p | 8 x 12.5 = 100 cents | $1.00 |
| 1080p | 8 x 18.75 = 150 cents | $1.50 |
| 4K | 8 x 37.5 = 300 cents | $3.00 |
Limits to remember
These apply to every request on this page, from the Video Router and Video generation docs:
- Length is 3 to 10 whole seconds in generation modes; an edit takes no duration.
- Aspect ratio is 16:9 or 9:16; an edit takes no aspect ratio.
- Native synced audio is always on, and
generate_audio: falseis rejected. - There is no
bitrate_mode, no reference audio and noseed. - Billing is the provider list times 1.25 per output second, reserved at submit and shown in
usage.cost.
Sources
Related posts
More in Developers
- Pinterest video aspect ratio text, quoted, and a master check
Pinterest's video spec says 'shorter than 1:2, taller than 1.91:1'. Quote it, convert the recommended ratios to decimals, and check a master with Sume frames.
- Pointing the OpenAI SDK at api.sume.com: why Sora calls still fail
Changing only the base URL does not carry a Sora call to Sume. Five documented differences: model ids, body format, seconds and size, status words, download.
- Poll a Sume image job in Python: obey next_poll_after_seconds
When POST /v1/images returns 202, keep polling the status_url and sleep for next_poll_after_seconds. A short Python loop that stops on a terminal status.
- Poll or callback for 1,000 video jobs: request counts at 30 seconds
At a 30-second poll, 1,000 video jobs make 4,000 to 20,000 status requests by job time. A callback_url cuts that to 1,000 deliveries plus a sweep.
Written by Sume