Gemini API's $10 per 10 minutes limit: how many Omni clips fit

Gemini API spend limits are $10, $50 or $200 per rolling 10 minutes by tier. At about $0.10 a second that is roughly 9 to 197 ten-second Omni clips per window.

5 min readSume
All posts

On the Gemini API, spend-based rate limits cap how many dollars you can spend in a rolling 10-minute window: $10 on Tier 1, $50 on Tier 2 and $200 on Tier 3. Google's pricing page puts Gemini Omni Flash 1.1 at about $0.10 per second of 720p video, so a ten-second 720p clip is roughly $1.01 and a Tier 1 project fits about nine of them per window. Sume paces differently, by plan concurrency and a queue rather than a dollar window.

What Google's page actually says

Google's rate limits page lists spend-based rate limits evaluated on a rolling 10-minute window, at $10 for Tier 1, $50 for Tier 2 and $200 for Tier 3.

Google's pricing page says Omni Flash video output is counted at 5,792 tokens per second of 720p video, about $0.10 per second at the $17.50 per million token standard rate, so exactly $0.1014 per second. Whether a given job is charged against the window at submit time or at completion is not stated on either page, so treat the table below as a planning ceiling and test with your own project.

Clips per window at the 720p list rate

Clips per 10-minute spend window at 5,792 tokens per second and $17.50 per million tokens (Google pages, read 2026-10-04)
TierSpend limit per 10 minutes10 s clips (about $1.01)8 s clips (about $0.81)
Tier 1$10912
Tier 2$504961
Tier 3$200197246

The same batch on Sume

Sume's Video Router bills provider list times 1.25 per output second, so the same ten-second 720p Omni clip is about $1.25 on Sume. The pacing controls are different: per Generation admission, a Pro workspace processes 4 jobs at once and can hold 20 more as queued, a Startup workspace 8 and 40, Scale 20 and 100. Submitting more than the processing cap is not an error; extra valid jobs wait in queued. When the queue is full you get 429 queue_full, and a spendable-balance shortfall returns 402 insufficient_credits before any provider work starts.

Sume documents no dollar-per-10-minutes window. Your real throughput is bounded by how fast jobs finish, which the docs do not quantify, so size by concurrency, not by a dollar figure.

Sizing a batch either way

Compute the ceiling in whole micro-dollars so rounding never lets one extra clip through. This snippet prints how many clips of a given length fit in each Google window.

Then submit to Sume with an Idempotency-Key per clip and poll with the interval the status response suggests, as described in Jobs and results.

# 5,792 tokens per second of 720p video at $17.50 per million tokens
MICROS_PER_SECOND = 5792 * 17_500_000 // 1_000_000  # 101,360
WINDOW_MICROS = {"tier1": 10_000_000, "tier2": 50_000_000, "tier3": 200_000_000}

def clips_per_window(seconds: int) -> dict:
    clip = MICROS_PER_SECOND * seconds
    return {tier: limit // clip for tier, limit in WINDOW_MICROS.items()}

for seconds in (8, 10):
    print(seconds, clips_per_window(seconds))

Which to pick

If you run a handful of clips, either route works. If you run hundreds, the Gemini API makes you manage a dollar window and a tier upgrade path; Sume makes you manage plan concurrency and the 25 percent premium over list. Neither page gives a time-to-finish guarantee, so measure ten real jobs before you promise a client a delivery date.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume