Google Vids makes free Omni 1080p scenes: when to use an API
Google Vids added Gemini Omni 1.1 Flash on Sept 23: 1080p scenes, extend and clip duration, free with a Google account. Where Sume's API route fits instead.

Google Vids now generates video scenes with Gemini Omni 1.1 Flash at no cost for anyone with a Google or Google Workspace account, per Google's September 23 post. If you need one clip for a deck, use Vids. If you need many clips from a script, a fixed size, or a file that lands in your own storage, call the model through an API such as Sume's.
What did Google add to Vids?
Google's post says anyone with a Google or Google Workspace account can generate high-quality videos at no cost with Gemini Omni 1.1 Flash, from vids.new on desktop. It lists 1080p HD scenes, extending scenes with smooth transitions, setting specific clip durations, upscaling existing AI clips, ready-made templates and a SynthID digital watermark in the video frames. It also says Workspace Business and Enterprise plans get expanded generation pools, and that text-to-speech narration in more than 100 languages, via Gemini 3.8 Flash-Lite, is coming soon.
The post gives no daily clip count and no regional list, so I cannot tell you how many scenes a free account makes. Treat "no cost" as the plan price, not a promise of unlimited volume.
Where does an API still win?
Vids is an editor. It does not give a script a request id, a retry key, or a webhook. The Gemini API page lists Omni Flash as paid-tier only with no free tier, at $17.50 per million video output tokens, which Google's pricing page converts at 5,792 tokens per second of 720p video, about $0.10 a second.
Sume lists the same model as gemini-omni-flash-1.1 and bills provider list times 1.25 per output second by resolution, so 720p is roughly $0.125 a second before any rounding. What you buy for that margin is the job envelope: an idempotency key, a status URL, a durable result URL and a catalog you can switch models in.
| Question | Google Vids | Gemini API or Sume |
|---|---|---|
| Cost | No cost with a Google account; expanded pools on Workspace plans | Paid per output second; no Gemini API free tier |
| Resolution | 1080p scenes per Google's post | 360p, 720p, 1080p, 4K (the Gemini API page says 1080p and 4K are upscaled) |
| Batch from a script | Not described | One request per clip, retry-safe with a key |
| Clip length | Set by you in the editor | 3 to 10 seconds per generation |
| Provenance | SynthID in the frames | SynthID applies to Omni output per Google |
What does the API call look like for the same scene?
On Sume, an Omni text-to-video job is one request. The Video Router docs list 3 to 10 seconds, 360p to 4K, and 16:9 or 9:16, with native audio always on.
Keep the prompt concrete and tell the model the camera and lighting, as you would in a Vids prompt box. The result comes back as a job you poll, as described in Video generation.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: scene-intro-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "A slow push-in on a whiteboard as a team lead points to a roadmap, soft office light",
"resolution": "1080p",
"duration": 8,
"aspect_ratio": "16:9",
"mode": "async"
}'What does Sume not give you here?
Sume has no editor timeline UI inside this call, no Vids templates, and no narration step bundled in. Vids' Workspace admin controls are also outside what Sume offers. Sume's docs say nothing about extending a Vids clip.
Pick by volume: one to five scenes for a presentation, Vids. A repeatable pipeline, Sume or the Gemini API directly. The comparison of Google surfaces in which Google surface for client work covers the account side.
Sources
Related posts
More in Comparisons
- GPT Image 2.5 high vs GPT Image 2 high: price per image
At 1024x1024 high, fal lists GPT Image 2 at $0.211 and GPT Image 2.5 at $0.05268, about a quarter. On Sume that is $0.264 vs $0.066 at list x 1.25.
- GPT Image 2.5 vs Gemini 3.1 Flash Image: price per 1K image
fal lists GPT Image 2.5 at $0.05268 per 1024x1024 high image; Google lists Gemini 3.1 Flash Image at about $0.067 per 1K image.
- grok-voice-transcribe-1.0 ended Oct 2: what a silent reroute means
xAI ended grok-voice-transcribe-1.0 on Oct 2, 2026 and routes it to 2.0 at the same price. How to catch silent model swaps, and what Sume STT 1.0 fixes for you.
- Ideogram 4.5 Magic Fill and Extend vs Sume mask_url and aspect ratio
Ideogram's docs list Magic Fill and Extend as editing features. Sume has no Extend endpoint; here is what its mask and aspect-ratio fields cover instead.
Written by Sume