Polling Omni jobs can't 429 your submits on Sume
Sume gives each API key separate read and write budgets: 120 writes and 4,800 reads a minute on Free, 1,200 and 48,000 on Scale. The math for an Omni batch.

On Sume, polling the status of Gemini Omni Flash jobs does not use the budget your submits use. Each API key has a requests-per-minute budget for /v1, and reads and writes are separate buckets. A GET poll is a read; creating a run is a write. Reads get forty times the write number, so a tight poll loop cannot cause a 429 on your own submits. Plan table and rules read from Sume's authentication docs on 2026-10-08.
Google's rate-limits page works differently and, as read on the same day, lists no Veo or Omni rows: it describes RPM, TPM and RPD per project and points to AI Studio for your active limits.
The budgets
Writes are the number the plan sells. Reads are the same number times 40.
| Plan | Writes | Reads |
|---|---|---|
| Free | 120 | 4,800 |
| Pro | 300 | 12,000 |
| Startup | 600 | 24,000 |
| Scale | 1,200 | 48,000 |
| Enterprise | Contact sales (Scale row until provisioned) | Contact sales |
Math for 100 Omni jobs
Say you have 100 Omni jobs in flight and you poll each every 5 seconds. That is 100 times 12 polls a minute, or 1,200 reads a minute, a quarter of the Free read budget. Submitting the 100 jobs costs 100 writes, which is under the Free budget of 120 in one minute. Request rate is not generation capacity, though: Free processes 1 job at a time and Pro processes 4, whatever your request rate is.
A webhook per job removes polling entirely. Sume signs the raw JSON body and sends x-sume-webhook-timestamp and x-sume-webhook-signature headers.
Reading a 429
Do not count requests yourself. Each response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, and a 429 adds retry-after. The error names the bucket in error.details.scope, either read or write. A 429 queue_full is a different failure: the workspace has no accepted-job room left, see Generation admission.
An MCP tool call spends the write budget once, for the run it creates. A jobs_status poll over MCP spends none.
Cost of being wrong
The expensive mistake is a retry without an idempotency key after a timeout: it can start a second paid job. At $1.00 for an 8-second 720p clip, 100 duplicates cost $100.00. Send a stable key, and a replay returns the original job.
A polling loop that waits
Poll every few seconds with a growing delay, and read retry-after if a 429 ever arrives. Because Pro allows 12,000 reads per minute against 300 writes, one key can poll 40 times as often as it submits. The number to watch is the accepted-job count, not the read rate.
Sources
Related posts
More in Developers
- Port an OpenRouter video client to Sume: base URL, key, model ids
Sume's /v1/videos follows the OpenRouter video wire. Three edits move a client: base URL, API key, bare model id. The six differences that still bite.
- PowerShell: Invoke-RestMethod for a 30-second Wan 3.0 clip
Windows PowerShell script that posts a 30-second wan-3.0 job, loops until it completes and saves the MP4 with Invoke-WebRequest. $3.75 at 720p.
- Pre-flight an Omni Flash request against the Sume model catalog
A short Python check that reads supported durations, resolutions and aspect ratios for gemini-omni-flash-1.1 via /v1/videos/models.
- Preflight a 3-minute Short with Timeline plan before you pay
POST /v1/timeline-1.0/plan compiles a Short without a job or a reserve and returns billable minutes and the estimate. A runnable Python check for six slots.
Written by Sume