Gemini Omni rate limit: Google lists none, Sume lists them
Google's rate-limit page has no Omni or Veo row. It lists per-project RPM, TPM and RPD plus spend caps. Sume publishes per-key limits and concurrency by plan.

Google's Gemini API rate-limits page, read 2026-10-08, has no row for Gemini Omni or Veo. It describes limits in requests per minute (RPM), tokens per minute (TPM) and requests per day (RPD), applied per project, and it caps spend per 10 minutes by tier: $10 on Tier 1, $50 on Tier 2 and $200 on Tier 3. For model-specific video limits it points you to AI Studio. If you need a number you can plan around, Sume publishes one for its own API.
So the honest answer to the query is that Google does not publish an Omni limit on that page.
What each side publishes
The Google row set is about projects. Sume's is about API keys and plans. They measure different things, so compare the structure, not the numbers.
| Item | Google Gemini API | Sume API |
|---|---|---|
| Omni or Veo row on the limits page | None | Not applicable; one limit for all models |
| Unit | Per project: RPM, TPM, RPD | Per key per minute: writes and reads |
| Spend cap | $10 / $50 / $200 per 10 min (Tier 1 / 2 / 3) | Credit balance |
| Concurrent jobs | Not stated on the page | Free 1, Pro 4, Startup 8, Scale 20 |
| Writes per minute | Not stated | Free 120, Pro 300, Startup 600, Scale 1,200 |
| Reads per minute | Not stated | Free 4,800, Pro 12,000, Startup 24,000, Scale 48,000 |
Sume's limits in practice
Sume reads are allowed at 40 times the write rate, so polling a job does not use up your submit budget. A 429 names error.details.scope as read or write, and responses carry ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after headers. A separate admission control limits queued and accepted jobs: Pro allows 4 running, 20 queued and 24 accepted, and a request beyond that returns 429 queue_full. A job you cannot afford returns 402 insufficient_credits.
What to do with an unknown limit
If you call Google directly, open AI Studio for the live limit for your project and treat the number as something that can change. Build a retry that honors the server's delay. If you call Sume, read retry-after and use the scope in the error to decide whether to slow submits or polling. Either way, submit in waves sized to the concurrency, not to the request rate.
Reading a 429 from Sume
A rate-limit 429 from Sume carries error.details.scope, which is read or write. A queue_full 429 is a different case: you are within the request rate but over the accepted-job count for your plan (24 on Pro, 48 on Startup, 120 on Scale, 6 on Free). Wait for jobs to finish; slowing the request rate will not help.
A 402 insufficient_credits means the balance cannot cover the reserve at submit. Add credit and resubmit with the same Idempotency-Key so you are not billed twice for one clip.
Sources
Related posts
More in Models
- GLM-5.3 reasoning cannot be turned off: cap the Sume run instead
Z.ai says GLM-5.3 always reasons, with low, high and max levels. What that means for run time and for a spend cap on a Sume agent run.
- GLM 5.3 Fast or Flash: which one does Sume list?
Sume's agent model list has GLM 5.3 Flash, not GLM 5.3 Fast. What Z.ai's GLM-5.3 page says and what the API's model field accepts.
- Does Google keep Omni and Veo prompts for 55 days?
Google's Gemini API page says prompts, context and outputs are kept 55 days for abuse checks. Veo files last 2 days. How that differs from a Sume request.
- Higgsfield Soul on Sume: half a cent per image, batches of 1 or 4
higgsfield-soul costs $0.005 at 720p and $0.0075 at 1080p on Sume. It takes batch sizes 1 or 4, no references. Price table and a four-image call.
Written by Sume