Veo 3.1 vs Gemini Omni Flash 1.1: what Sume can call

Google lists Veo 3.1 at 1080p and 4K with native audio. On Sume, Gemini Omni Flash 1.1 is the callable Google id: 3-10 s, 360p-4K, audio always on.

4 min readSume
All posts

If you are choosing between Veo 3.1 and Gemini Omni Flash 1.1, the Sume answer is simple: gemini-omni-flash-1.1 is in the Sume Video Router catalog and Veo 3.1 is not an id you can call on Sume today. Omni Flash on Sume gives you 3 to 10 second clips at 360p, 720p, 1080p or 4K, in 16:9 or 9:16, with synced audio always on, plus a video edit mode that Veo's page does not describe.

Everything about Veo below comes from Google DeepMind's own Veo page, read on 2026-10-03. That page is a capability overview and sends readers to a model card for exact parameters, so this comparison leaves out any Veo limit it does not state.

Side by side

The table keeps the two columns honest: the Veo column is only what the DeepMind page says, and the Sume column is only what the Sume docs say about the catalog id.

Veo 3.1 (Google page) vs Gemini Omni Flash 1.1 (Sume docs), read 2026-10-03
ItemVeo 3.1 (Google DeepMind page)Gemini Omni Flash 1.1 on Sume
Resolution1080p and 4K options360p, 720p, 1080p, 4K
Length8 seconds mentioned as standard3-10 seconds
AudioNative audio: effects, ambience, dialogueNative synced audio, always on
ReferencesReference images for characters, scenes, objectsUp to 10 images and up to 3 videos of 3 s or less
EditingOutpainting to other screen shapesVideo-to-video edit via video_url
Callable on SumeNoYes, gemini-omni-flash-1.1

Where Omni Flash differs in practice

Audio is the first behavioral difference to plan for. On Sume, generate_audio: false is rejected for Omni Flash because audio is always on, and the model has no reference_audio_urls. If you need a silent clip you will mute it afterwards or choose another catalog model. Google's page says natural spoken audio for short speech segments is still an area of active development, which is worth testing on any short dialogue clip regardless of the model.

The second difference is request shape. One id covers four capabilities and Sume picks by what you send: prompt alone is text-to-video, image_url (with optional end_image_url) is image-to-video, reference_image_urls and reference_video_urls is reference-to-video, and video_url is an edit. Reference media is addressed in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>, counted from zero in list order. An edit source cannot be combined with image or reference fields.

Billing follows the Video Router rule: provider list price times 1.25 per output second, by resolution. Pull the live numbers from GET /v1/video-router/models instead of hard-coding them.

An edit request

``bash curl -X POST https://api.sume.com/v1/video-router/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: omni-edit-001" \ -d '{ "model": "gemini-omni-flash-1.1", "prompt": "Replace the bottle with an apple. Keep everything else the same.", "video_url": "https://example.com/clip.mp4", "resolution": "720p", "mode": "async" }' ``

For an edit, leave out aspect_ratio and duration; the docs say they are not accepted in that mode.

Which to use

Use Veo where Google offers it if the 8-second default and its own interface suit the job. Use Omni Flash on Sume when you need 10 seconds, 4K, native audio and an edit mode behind a single async job API. The related posts on Veo 3.1 versus Grok Imagine on Sume and what to do when you need more than 8 seconds go further on clip length.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume