Gemini API leads with Omni Flash, Veo 3.1 for specialists: and Sume?
Google's video docs recommend Gemini Omni Flash as the default and Veo 3.1 for specific needs. Sume has sume/auto or a pinned catalog id. How they line up.

Google's Gemini API video page now offers two models and recommends Gemini Omni Flash as the default, with Veo 3.1 for specialized needs. Sume offers a routing mode, sume/auto, and lists Gemini Omni Flash 1.1 as a pinnable catalog id. The catalogs differ, so check the live model list before you pin an id.
What Google's page says
On the page I read, Gemini Omni Flash is recommended as the default model for video generation, with text, image, audio and video inputs and multi-turn conversational editing. Veo 3.1 is described as generating video with native audio and supporting video extension, frame-specific generation and image-based direction; Google points to it for scene extension, last-frame control or legacy pipelines. The page excerpt I could read did not include limits or prices, so this post quotes none.
| Platform | Default | Specialized option |
|---|---|---|
| Gemini API | Gemini Omni Flash | Veo 3.1: native audio, extension, frame-specific generation |
| Sume /v1/videos | sume/auto, Sume selects the model | Pin any catalog id, such as seedance-2.5, wan-3.0 or gemini-omni-flash-1.1 |
How sume/auto behaves
Send model: "sume/auto" on POST /v1/videos and Sume selects the model. The poll response reports sume/auto and does not disclose which family ran. The selection is a pure function of the request and the catalog version, so the same request resolves the same way until the catalog changes.
When to pin a model instead
- You need a specific look that you validated on one model.
- You need a capability that only some models have, for example reference audio or a longer duration.
- You want a price you can predict per second for budgeting.
- You need the response to name the model for your own records.
Check the live list
Sume's catalog lives at GET /v1/videos/models, and the limits differ by model, so read the capabilities from the catalog instead of assuming one envelope. The ids named in the Sume video docs I read include seedance-2, seedance-2.5, wan-3.0, minimax-h3-max and gemini-omni-flash-1.1, and none is a Veo id. Anything the live catalog does not list is not available through Sume, whatever the docs say.
If your application depends on Veo 3.1 specifically, such as its video extension, call Google's API for that model. Sume does not claim to proxy it.
A sensible default
Start every new use case on sume/auto, record the model name only if you pin one, and move to a pinned id when a test shows a clear reason. Use the same ten test prompts each time so the comparison stays fair.
Sources
Related posts
More in Comparisons
- Grok Imagine, Qwen Image and Imagen 4 Fast all cost 2.5 cents on Sume
Grok Imagine, Qwen Image and Imagen 4 Fast each bill $0.025 per image on Sume. What separates them: ratios, n, edits. Choose the cheap row that fits your job.
- HappyHorse 1.0 vs 1.1: which one renders 480P on Model Studio?
On Alibaba Model Studio only HappyHorse 1.1 lists 480P. Both versions take 3 to 15 seconds, and 1.0 adds a video-edit id. Neither is in Sume's catalog.
- Headshot background to neutral grey: cutout plus Pillow or an AI edit
Swap a headshot background for grey: Sume RMBG at $0.0225 plus a Pillow composite keeps the face pixels exact; an Ideogram 4.5 edit costs $0.075 or more.
- HeyGen's X-RateLimit-Scope header vs Sume's error.details.scope
HeyGen's October 2026 429s add a scope header. Sume names the bucket in error.details.scope and sends ratelimit-* headers; one parser can read both.
Written by Sume