Gemini Omni video in the Windows app, and by API on Sume
Google lists Gemini Omni video in the Gemini app for Windows. For your own app, Sume exposes gemini-omni-flash-1.1 at 360p to 4K, 3 to 10 seconds.

If you want Gemini Omni video inside your own product rather than in a desktop app, call gemini-omni-flash-1.1 on Sume with POST /v1/videos. It takes 3 to 10 seconds at 360p, 720p, 1080p or 4K, in 16:9 or 9:16, and always returns audio.
Google's September 2026 roundup, read 2026-10-05, mentions generating Gemini Omni videos in the Gemini app for Windows. That is a consumer surface, and the page we read gives no API prices or limits, so we do not quote any from it.
App versus API
A desktop app is built for one person making one clip at a time. An API is for batches, queues and your own storage. On Sume the job is asynchronous: you submit, poll the status, then download from /content, and you can pass a callback_url over HTTPS so your server hears about completion.
The Sume row is also the default for sume/auto, so a request with no pinned model resolves to it (8 seconds, 720p, unless you set otherwise).
What the row takes
Limits and modes from the Video Router docs, prices as list times 1.25 rounded up per job. fal's list comes from its Omni Flash 1.1 page, read the same day.
| Resolution | fal list per second | Sume per second | Sume 10 s clip |
|---|---|---|---|
| 360p | $0.03 | $0.0375 | $0.38 |
| 720p | $0.10 | $0.125 | $1.25 |
| 1080p | $0.15 | $0.1875 | $1.88 |
| 4K | $0.30 | $0.375 | $3.75 |
Modes you cannot get from an app
- Image-to-video with a first frame and an optional last frame.
- Reference-to-video with up to 10 images and up to 3 clips of at most 3 seconds each, addressed in the prompt as
<IMAGE_REF_0>and<VIDEO_REF_0>. - Edit mode, on the Video Router
/v1/video-router/generatesurface: send avideo_urland a prompt, and the source clip sets the output length (noaspect_ratioallowed). - Idempotent retries with an
Idempotency-Keyheader.
Limits to plan around
The row has no audio toggle: generate_audio: false returns 400, because audio is always on. It rejects size, seed and bitrate_mode, and the longest clip is 10 seconds. For anything longer, split the scene or pick a model with a higher cap such as wan-3.0 (30 seconds). The 4K post covers when the top tier is worth the price.
Sources
Related posts
More in Models
- generate_audio false on Gemini Omni: accepted, clip keeps its audio
Video Router accepts generate_audio false on gemini-omni-flash-1.1 and ignores it; MiniMax H3 answers 400. Check the model before you rely on a mute flag.
- GPT Image 2.5 above 2560x1440 is experimental: a safe size ladder
OpenAI marks sizes over 2560x1440 experimental on GPT Image models. A ladder of valid sizes up to 3840x2160 and a Python check for Sume's image_size.
- GPT Image 2.5 quality: OpenAI defaults to auto, Sume to high
OpenAI's default quality for GPT Image 2.5 is auto; Sume's is high when omitted. Why that differs, what auto reserves on Sume, and how to pin quality.
- gpt-live-transcribe: realtime STT at $0.017 a minute
gpt-live-transcribe costs $0.017 a minute ($1.02 an hour) and runs only on the Realtime transcription sessions endpoint. Use gpt-transcribe for files.
Written by Sume