Omni Flash or Veo 3.1 for the Gemini API? What Sume runs
Google's video docs now call Omni Flash the default over Veo 3.1. Here are Google's listed prices for both, and which one Sume's video router carries.

Should you use Omni Flash or Veo 3.1 through the Gemini API? Google's video generation docs recommend Omni Flash "as the default choice for video generation" and still list Veo 3.1. Sume's video router carries Omni (gemini-omni-flash-1.1) and does not carry Veo 3.1, so for Sume callers the choice is already made.
Google's listed prices
Google's pricing page prices Omni gemini-omni-1.1-flash video output at $17.50 per million tokens, about $0.10 per second at 720p. Veo 3.1 is priced per second in three tiers.
| Model | 720p | 1080p | 4K |
|---|---|---|---|
| Omni 1.1 Flash (approx.) | $0.10 | not listed | not listed |
| Veo 3.1 Standard | $0.40 | $0.40 | $0.60 |
| Veo 3.1 Fast | $0.10 | $0.12 | $0.30 |
| Veo 3.1 Lite | $0.05 | $0.08 | no 4K |
What that means for choosing
At 720p Omni sits at the same price as Veo 3.1 Fast, and above Lite. The Omni docs add scene extension to 40 seconds, multi-turn editing with previous_interaction_id and video references up to 3 clips of 3 seconds each, which are features beyond a plain prompt-to-clip call.
Veo 3.1 Lite is the cheapest line on the page, so if you only need short silent-style drafts and call Google directly, it is worth a test.
On Sume
Sume lists Omni at 3 to 10 seconds, 360p to 4K, 16:9 and 9:16, with native audio. Billing is the list rate times 1.25, rounded up to cents, so a 10 second 720p clip comes to $1.25. Sume does not expose extension or previous_interaction_id, and its edit mode takes a video_url.
If a project needs Veo specifically, call Google for it; do not expect a Veo id from the Sume router.
Before you switch
A default recommendation in documentation describes where the vendor wants new users to start, not a ranking of output quality, and I have not benchmarked the two. Run your own prompt set through each, compare cost per usable second, not cost per generated second, and count retries. If a model returns two unusable clips for every good one, its sticker price understates what you pay.
Pricing and model lists on Google's page change often, so reread it on the day you commit.
Reference limits worth knowing
Omni takes video references of up to 3 clips, each up to 3 seconds, and payloads over 4 MB should use delivery="uri". On Sume the same row accepts up to 10 images and up to 3 reference videos, each at most 3 seconds, and the prompt may run up to 20,000 characters.
Sources
Related posts
More in Models
- Gemini Omni scene extension: app, Flow, API or Sume?
Google's launch post lists four places Omni 1.1 Flash runs. Scene extension is not in all of them, and Sume's router does not expose it. Where to go instead.
- gpt-image-1 retires December 1: scan your repo, move to gpt-image-2.5
OpenAI's deprecations page lists gpt-image-1, 1.5 and 1-mini for December 1, 2026. A Python scan finds the ids, and Sume serves openai/gpt-image-2.5.
- gpt-image-2 or gpt-image-2.5 to replace gpt-image-1 on Sume?
OpenAI names gpt-image-2.5 as the gpt-image-1 replacement. On Sume, gpt-image-2 and 2.5 differ on references, mask_url, background and quality tiers.
- Ideogram 4.5 headline variants from one ad: 4 edits for about $0.30
Four headline edits of one approved ad on ideogram/ideogram-v4.5 at medium quality cost about $0.30 by Sume's catalog line. Python that runs them in parallel.
Written by Sume