Google made Omni Flash its default video model: pin it or sume/auto?

Google's docs now say to use Gemini Omni Flash as the default video model. On Sume you can pin gemini-omni-flash-1.1 or send sume/auto. What each one fixes.

4 min readSume
All posts

Pin gemini-omni-flash-1.1 when your product depends on what Omni does: 3 to 10 seconds, native audio, up to 10 image references and 3 short video references, and the edit mode. Send sume/auto when you want Sume to choose a model per request and you can accept that the response will not say which one ran. Google's own docs now say to use Gemini Omni Flash as the default for video generation, which is a vendor recommendation, not a Sume default.

What Google changed

Google's video generation page lists two models, Gemini Omni Flash and Veo 3.1, and tells readers to use Omni Flash as their default model. Its deprecations page lists the three Veo 3.1 preview ids with an October 22, 2026 shutdown date and gemini-omni-1.1-flash as the replacement. That is a change of default on Google's side. It does not change what Sume lists, and Sume does not list a Veo model.

Pinning versus sume/auto

The Sume docs describe both routes on the same POST /v1/videos endpoint. A pinned request names a catalog model and gets that model's limits. An sume/auto request lets Sume select the family, and the poll response reports sume/auto; Sume does not disclose which family ran and tells you not to infer it from the output.

Pinned Omni versus sume/auto on /v1/videos, from Sume docs read 2026-10-08
Questionmodel: gemini-omni-flash-1.1model: sume/auto
Who chooses the modelYouSume
Clip length3 to 10 seconds3 to 10 seconds in the Auto controls
Default resolution and lengthYou choose; resolution 360p to 4K720p and 8 seconds
Aspect ratios16:9 or 9:1616:9 or 9:16
Model named in the responsegemini-omni-flash-1.1sume/auto
Replay with the same Idempotency-KeySame job returnedSame price and same route

What the table does not decide

Neither choice changes what you pay per second for Omni: both routes bill through your workspace balance, with the provider list times 1.25 as the rule for pinned Omni. With Auto the route can differ between requests, so the price can too. The safest way to compare is to submit the same brief both ways with the same duration and resolution and read usage.cost from each poll response. Sume's docs say the Auto create controls default to 720p and 8 seconds, so an Auto request with no overrides is a different shape from a pinned request where you picked 1080p.

When pinning is the right call

Pin when a client has approved the look of one model, when you reuse its reference-media syntax (<IMAGE_REF_0> and <VIDEO_REF_0>, 0-based, in list order), when you need the 4K row, or when you bill your own customers a price you computed from one per-second rate. A pinned id also makes your cost predictable: 8 seconds at 720p is 8 x $0.125 = $1.00 on every run.

When sume/auto is the right call

Use Auto when the brief is a vertical product clip and nobody cares which model draws it, when you want your integration to survive a model retirement without a code change, or when you are prototyping. Read usage.cost from each job because the price depends on the route Sume chose.

A simple policy that works for many teams: pin Omni for finals that were approved on Omni, use Auto for drafts, and keep the model id in configuration so that a vendor's next change is an edit to one value, not a deploy.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume