Google made Omni Flash its default video model: pin it or sume/auto?
Google's docs now say to use Gemini Omni Flash as the default video model. On Sume you can pin gemini-omni-flash-1.1 or send sume/auto. What each one fixes.

Pin gemini-omni-flash-1.1 when your product depends on what Omni does: 3 to 10 seconds, native audio, up to 10 image references and 3 short video references, and the edit mode. Send sume/auto when you want Sume to choose a model per request and you can accept that the response will not say which one ran. Google's own docs now say to use Gemini Omni Flash as the default for video generation, which is a vendor recommendation, not a Sume default.
What Google changed
Google's video generation page lists two models, Gemini Omni Flash and Veo 3.1, and tells readers to use Omni Flash as their default model. Its deprecations page lists the three Veo 3.1 preview ids with an October 22, 2026 shutdown date and gemini-omni-1.1-flash as the replacement. That is a change of default on Google's side. It does not change what Sume lists, and Sume does not list a Veo model.
Pinning versus sume/auto
The Sume docs describe both routes on the same POST /v1/videos endpoint. A pinned request names a catalog model and gets that model's limits. An sume/auto request lets Sume select the family, and the poll response reports sume/auto; Sume does not disclose which family ran and tells you not to infer it from the output.
| Question | model: gemini-omni-flash-1.1 | model: sume/auto |
|---|---|---|
| Who chooses the model | You | Sume |
| Clip length | 3 to 10 seconds | 3 to 10 seconds in the Auto controls |
| Default resolution and length | You choose; resolution 360p to 4K | 720p and 8 seconds |
| Aspect ratios | 16:9 or 9:16 | 16:9 or 9:16 |
| Model named in the response | gemini-omni-flash-1.1 | sume/auto |
| Replay with the same Idempotency-Key | Same job returned | Same price and same route |
What the table does not decide
Neither choice changes what you pay per second for Omni: both routes bill through your workspace balance, with the provider list times 1.25 as the rule for pinned Omni. With Auto the route can differ between requests, so the price can too. The safest way to compare is to submit the same brief both ways with the same duration and resolution and read usage.cost from each poll response. Sume's docs say the Auto create controls default to 720p and 8 seconds, so an Auto request with no overrides is a different shape from a pinned request where you picked 1080p.
When pinning is the right call
Pin when a client has approved the look of one model, when you reuse its reference-media syntax (<IMAGE_REF_0> and <VIDEO_REF_0>, 0-based, in list order), when you need the 4K row, or when you bill your own customers a price you computed from one per-second rate. A pinned id also makes your cost predictable: 8 seconds at 720p is 8 x $0.125 = $1.00 on every run.
When sume/auto is the right call
Use Auto when the brief is a vertical product clip and nobody cares which model draws it, when you want your integration to survive a model retirement without a code change, or when you are prototyping. Read usage.cost from each job because the price depends on the route Sume chose.
A simple policy that works for many teams: pin Omni for finals that were approved on Omni, use Auto for drafts, and keep the model id in configuration so that a vendor's next change is an edit to one value, not a deploy.
Sources
Related posts
More in Developers
- Google says Imagen is shut down: test your Imagen call on Sume now
Google's Imagen page says Imagen models are shut down. Sume lists Imagen 4 Fast and Ultra separately. One catalog check and one request tell you if yours works.
- GPT Image 2.5 inpainting with mask_url on Sume: steps and cost
Edit part of an image with mask_url and up to 16 references on openai/gpt-image-2.5. Billed about $0.066 an image on Sume. Steps and the limits.
- gpt-image-2.5 quality auto reserves max: holds from $0.22 to $0.89
On Sume, gpt-image-2.5 with quality auto and auto size reserves $0.8895 per image, while omitting quality reserves $0.2224. Hold table and the safe request.
- Grok Image n is 1 on Sume: four images mean four calls, $0.10
Sume's catalog caps n at 1 for Grok Image. A Python loop for four images costs $0.10, and runs sync. Ratios incl. 9:19.5 and 9:20 stay available.
Written by Sume