Omni Flash or Veo 3.1 for the Gemini API? What Sume runs

Google's video docs now call Omni Flash the default over Veo 3.1. Here are Google's listed prices for both, and which one Sume's video router carries.

4 min readSume
All posts

Should you use Omni Flash or Veo 3.1 through the Gemini API? Google's video generation docs recommend Omni Flash "as the default choice for video generation" and still list Veo 3.1. Sume's video router carries Omni (gemini-omni-flash-1.1) and does not carry Veo 3.1, so for Sume callers the choice is already made.

Google's listed prices

Google's pricing page prices Omni gemini-omni-1.1-flash video output at $17.50 per million tokens, about $0.10 per second at 720p. Veo 3.1 is priced per second in three tiers.

Gemini API video prices per second, read 2026-10-06
Model720p1080p4K
Omni 1.1 Flash (approx.)$0.10not listednot listed
Veo 3.1 Standard$0.40$0.40$0.60
Veo 3.1 Fast$0.10$0.12$0.30
Veo 3.1 Lite$0.05$0.08no 4K

What that means for choosing

At 720p Omni sits at the same price as Veo 3.1 Fast, and above Lite. The Omni docs add scene extension to 40 seconds, multi-turn editing with previous_interaction_id and video references up to 3 clips of 3 seconds each, which are features beyond a plain prompt-to-clip call.

Veo 3.1 Lite is the cheapest line on the page, so if you only need short silent-style drafts and call Google directly, it is worth a test.

On Sume

Sume lists Omni at 3 to 10 seconds, 360p to 4K, 16:9 and 9:16, with native audio. Billing is the list rate times 1.25, rounded up to cents, so a 10 second 720p clip comes to $1.25. Sume does not expose extension or previous_interaction_id, and its edit mode takes a video_url.

If a project needs Veo specifically, call Google for it; do not expect a Veo id from the Sume router.

Before you switch

A default recommendation in documentation describes where the vendor wants new users to start, not a ranking of output quality, and I have not benchmarked the two. Run your own prompt set through each, compare cost per usable second, not cost per generated second, and count retries. If a model returns two unusable clips for every good one, its sticker price understates what you pay.

Pricing and model lists on Google's page change often, so reread it on the day you commit.

Reference limits worth knowing

Omni takes video references of up to 3 clips, each up to 3 seconds, and payloads over 4 MB should use delivery="uri". On Sume the same row accepts up to 10 images and up to 3 reference videos, each at most 3 seconds, and the prompt may run up to 20,000 characters.

Sources

Related posts

More in Models

All Models posts

Written by Sume