Veo 3.1 vs Gemini Omni Flash 1.1: what Sume can call
Google lists Veo 3.1 at 1080p and 4K with native audio. On Sume, Gemini Omni Flash 1.1 is the callable Google id: 3-10 s, 360p-4K, audio always on.

If you are choosing between Veo 3.1 and Gemini Omni Flash 1.1, the Sume answer is simple: gemini-omni-flash-1.1 is in the Sume Video Router catalog and Veo 3.1 is not an id you can call on Sume today. Omni Flash on Sume gives you 3 to 10 second clips at 360p, 720p, 1080p or 4K, in 16:9 or 9:16, with synced audio always on, plus a video edit mode that Veo's page does not describe.
Everything about Veo below comes from Google DeepMind's own Veo page, read on 2026-10-03. That page is a capability overview and sends readers to a model card for exact parameters, so this comparison leaves out any Veo limit it does not state.
Side by side
The table keeps the two columns honest: the Veo column is only what the DeepMind page says, and the Sume column is only what the Sume docs say about the catalog id.
| Item | Veo 3.1 (Google DeepMind page) | Gemini Omni Flash 1.1 on Sume |
|---|---|---|
| Resolution | 1080p and 4K options | 360p, 720p, 1080p, 4K |
| Length | 8 seconds mentioned as standard | 3-10 seconds |
| Audio | Native audio: effects, ambience, dialogue | Native synced audio, always on |
| References | Reference images for characters, scenes, objects | Up to 10 images and up to 3 videos of 3 s or less |
| Editing | Outpainting to other screen shapes | Video-to-video edit via video_url |
| Callable on Sume | No | Yes, gemini-omni-flash-1.1 |
Where Omni Flash differs in practice
Audio is the first behavioral difference to plan for. On Sume, generate_audio: false is rejected for Omni Flash because audio is always on, and the model has no reference_audio_urls. If you need a silent clip you will mute it afterwards or choose another catalog model. Google's page says natural spoken audio for short speech segments is still an area of active development, which is worth testing on any short dialogue clip regardless of the model.
The second difference is request shape. One id covers four capabilities and Sume picks by what you send: prompt alone is text-to-video, image_url (with optional end_image_url) is image-to-video, reference_image_urls and reference_video_urls is reference-to-video, and video_url is an edit. Reference media is addressed in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>, counted from zero in list order. An edit source cannot be combined with image or reference fields.
Billing follows the Video Router rule: provider list price times 1.25 per output second, by resolution. Pull the live numbers from GET /v1/video-router/models instead of hard-coding them.
An edit request
``bash
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-edit-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Replace the bottle with an apple. Keep everything else the same.",
"video_url": "https://example.com/clip.mp4",
"resolution": "720p",
"mode": "async"
}'
``
For an edit, leave out aspect_ratio and duration; the docs say they are not accepted in that mode.
Which to use
Use Veo where Google offers it if the 8-second default and its own interface suit the job. Use Omni Flash on Sume when you need 10 seconds, 4K, native audio and an edit mode behind a single async job API. The related posts on Veo 3.1 versus Grok Imagine on Sume and what to do when you need more than 8 seconds go further on clip length.
Sources
Related posts
More in Comparisons
- Voice agent latency claims: 100 ms, 150 ms, 90 ms and Sume TTS jobs
ElevenLabs and Cartesia quote 90 to 150 ms for live voice. What those figures measure, and why Sume TTS jobs are for produced audio rather than live calls.
- Voice agent platforms at 10,000 minutes: eight price lists compared
Telnyx about $596, Deepgram and AssemblyAI $750, Bland $1,400 to $1,499, Retell $700 to $3,100: eight published rates at 10,000 minutes.
- Voice isolator vs audio detach: 500 MB and 1 hour vs 1800 seconds
ElevenLabs Voice Isolator removes background noise from speech. Sume audio detach extracts the video audio track as is; it does not separate vocals from music.
- Wan 3.0 or MiniMax H3 Max: which to pin for reference-to-video
Both take image, video and audio references on Sume. Wan runs 2 to 30 seconds with 5 reference videos; H3 Max runs 5 to 15 seconds with always-on stereo audio.
Written by Sume