Gemini API's $10 per 10 minutes limit: how many Omni clips fit
Gemini API spend limits are $10, $50 or $200 per rolling 10 minutes by tier. At about $0.10 a second that is roughly 9 to 197 ten-second Omni clips per window.

On the Gemini API, spend-based rate limits cap how many dollars you can spend in a rolling 10-minute window: $10 on Tier 1, $50 on Tier 2 and $200 on Tier 3. Google's pricing page puts Gemini Omni Flash 1.1 at about $0.10 per second of 720p video, so a ten-second 720p clip is roughly $1.01 and a Tier 1 project fits about nine of them per window. Sume paces differently, by plan concurrency and a queue rather than a dollar window.
What Google's page actually says
Google's rate limits page lists spend-based rate limits evaluated on a rolling 10-minute window, at $10 for Tier 1, $50 for Tier 2 and $200 for Tier 3.
Google's pricing page says Omni Flash video output is counted at 5,792 tokens per second of 720p video, about $0.10 per second at the $17.50 per million token standard rate, so exactly $0.1014 per second. Whether a given job is charged against the window at submit time or at completion is not stated on either page, so treat the table below as a planning ceiling and test with your own project.
Clips per window at the 720p list rate
| Tier | Spend limit per 10 minutes | 10 s clips (about $1.01) | 8 s clips (about $0.81) |
|---|---|---|---|
| Tier 1 | $10 | 9 | 12 |
| Tier 2 | $50 | 49 | 61 |
| Tier 3 | $200 | 197 | 246 |
The same batch on Sume
Sume's Video Router bills provider list times 1.25 per output second, so the same ten-second 720p Omni clip is about $1.25 on Sume. The pacing controls are different: per Generation admission, a Pro workspace processes 4 jobs at once and can hold 20 more as queued, a Startup workspace 8 and 40, Scale 20 and 100. Submitting more than the processing cap is not an error; extra valid jobs wait in queued. When the queue is full you get 429 queue_full, and a spendable-balance shortfall returns 402 insufficient_credits before any provider work starts.
Sume documents no dollar-per-10-minutes window. Your real throughput is bounded by how fast jobs finish, which the docs do not quantify, so size by concurrency, not by a dollar figure.
Sizing a batch either way
Compute the ceiling in whole micro-dollars so rounding never lets one extra clip through. This snippet prints how many clips of a given length fit in each Google window.
Then submit to Sume with an Idempotency-Key per clip and poll with the interval the status response suggests, as described in Jobs and results.
# 5,792 tokens per second of 720p video at $17.50 per million tokens
MICROS_PER_SECOND = 5792 * 17_500_000 // 1_000_000 # 101,360
WINDOW_MICROS = {"tier1": 10_000_000, "tier2": 50_000_000, "tier3": 200_000_000}
def clips_per_window(seconds: int) -> dict:
clip = MICROS_PER_SECOND * seconds
return {tier: limit // clip for tier, limit in WINDOW_MICROS.items()}
for seconds in (8, 10):
print(seconds, clips_per_window(seconds))Which to pick
If you run a handful of clips, either route works. If you run hundreds, the Gemini API makes you manage a dollar window and a tier upgrade path; Sume makes you manage plan concurrency and the 25 percent premium over list. Neither page gives a time-to-finish guarantee, so measure ten real jobs before you promise a client a delivery date.
Sources
- Google AI for Developers: Rate limits (read 2026-10-04)
- Google AI for Developers: Gemini Developer API pricing (read 2026-10-04)
- Google AI for Developers: Generate and edit videos with Gemini Omni Flash (read 2026-10-04)
- Generation admission (Sume docs)
- Video Router (Sume docs)
- Jobs and results (Sume docs)
Related posts
More in Developers
- Gemini CLI 0.62 MCP titles: reading Sume's tool names
Gemini CLI v0.62.0 formats MCP tool call titles as structured signatures. Sume tool ids are underscore names such as generate_image; dotted aliases map to them.
- Gemini CLI v0.63 plan execution in CI: gate paid Sume calls first
Gemini CLI preview v0.63.0 adds autonomous plan execution in non-interactive mode. Before unattended runs, gate Sume paid tools with dry_run and max_spend_usd.
- Headless Gemini CLI and MCP auth: use a Sume API key
A headless Gemini CLI run cannot finish a browser consent. Connect to hosted Sume MCP with an API key header instead, and keep paid calls bounded.
- Gemini CLI untrusted tool output provenance: Sume results as data
Gemini CLI 0.60 and 0.61 release notes name fixes for untrusted tool output and indirect prompt injection. Keep Sume write tools behind scope and caps.
Written by Sume