Gemini Omni rate limit: Google lists none, Sume lists them

Google's rate-limit page has no Omni or Veo row. It lists per-project RPM, TPM and RPD plus spend caps. Sume publishes per-key limits and concurrency by plan.

6 min readSume
All posts

Google's Gemini API rate-limits page, read 2026-10-08, has no row for Gemini Omni or Veo. It describes limits in requests per minute (RPM), tokens per minute (TPM) and requests per day (RPD), applied per project, and it caps spend per 10 minutes by tier: $10 on Tier 1, $50 on Tier 2 and $200 on Tier 3. For model-specific video limits it points you to AI Studio. If you need a number you can plan around, Sume publishes one for its own API.

So the honest answer to the query is that Google does not publish an Omni limit on that page.

What each side publishes

The Google row set is about projects. Sume's is about API keys and plans. They measure different things, so compare the structure, not the numbers.

Published limits, Google read 2026-10-08, Sume from its docs
ItemGoogle Gemini APISume API
Omni or Veo row on the limits pageNoneNot applicable; one limit for all models
UnitPer project: RPM, TPM, RPDPer key per minute: writes and reads
Spend cap$10 / $50 / $200 per 10 min (Tier 1 / 2 / 3)Credit balance
Concurrent jobsNot stated on the pageFree 1, Pro 4, Startup 8, Scale 20
Writes per minuteNot statedFree 120, Pro 300, Startup 600, Scale 1,200
Reads per minuteNot statedFree 4,800, Pro 12,000, Startup 24,000, Scale 48,000

Sume's limits in practice

Sume reads are allowed at 40 times the write rate, so polling a job does not use up your submit budget. A 429 names error.details.scope as read or write, and responses carry ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after headers. A separate admission control limits queued and accepted jobs: Pro allows 4 running, 20 queued and 24 accepted, and a request beyond that returns 429 queue_full. A job you cannot afford returns 402 insufficient_credits.

What to do with an unknown limit

If you call Google directly, open AI Studio for the live limit for your project and treat the number as something that can change. Build a retry that honors the server's delay. If you call Sume, read retry-after and use the scope in the error to decide whether to slow submits or polling. Either way, submit in waves sized to the concurrency, not to the request rate.

Reading a 429 from Sume

A rate-limit 429 from Sume carries error.details.scope, which is read or write. A queue_full 429 is a different case: you are within the request rate but over the accepted-job count for your plan (24 on Pro, 48 on Startup, 120 on Scale, 6 on Free). Wait for jobs to finish; slowing the request rate will not help.

A 402 insufficient_credits means the balance cannot cover the reserve at submit. Add credit and resubmit with the same Idempotency-Key so you are not billed twice for one clip.

Sources

Related posts

More in Models

All Models posts

Written by Sume