Sume submit budgets: 120 to 1,200 writes a minute for an Omni batch
Sume rate limits submits per plan: Free 120, Pro 300, Startup 600 and Scale 1,200 a minute; reads get 40 times that. Why queue size, not the rate, paces Omni.

Sume budgets submit calls per plan: Free 120 per minute, Pro 300, Startup 600 and Scale 1,200, and reads get 40 times the write budget. For Omni video these rates almost never bind, because the queue fills first: a Pro workspace accepts 24 jobs, then returns 429 queue_full. Exceeding the call rate itself returns 429 rate_limited, a different error with a different fix.
Two limits, two errors
Generation admission separates them. The submit budget protects the API; the queue protects the generation workers. A rate_limited response means you are sending too many calls per minute, so slow down. A queue_full response means you already hold the maximum accepted jobs, so wait for completions.
On Google's side, the rate limits page uses a third kind of limit, spend over a rolling 10 minutes, with no Omni request-per-minute figure.
| Plan | Submit writes per minute | Reads per minute | Accepted generation jobs |
|---|---|---|---|
| Free | 120 | 4,800 | 6 |
| Pro | 300 | 12,000 | 24 |
| Startup | 600 | 24,000 | 48 |
| Scale | 1,200 | 48,000 | 120 |
What binds first
A script that submits 30 Omni clips to a Pro workspace in one second is at 30 calls, far below 300 per minute, but it will get 6 queue-full responses once 24 are accepted. So pace by accepted jobs, not by calls.
Use the wave_size_hint the docs describe, the larger of 1 and 75 percent of remaining queue capacity, to size each wave. Poll the status endpoint at the suggested interval; reads are cheap but not free, and the status response includes next_poll_after_seconds, as described in Jobs and results.
Cost of the same wave
At the Video Router rate of list times 1.25, a full Pro queue of 24 ten-second 720p clips comes to $30.00 at about $0.10 per second list times 1.25, and Sume reserves the cost at submit. The same 24 clips at Google's list rate would be about $24, but 24 clips in one window exceeds the $10 Tier 1 spend limit, so a Tier 1 project would need three or more windows.
Sources
Related posts
More in Developers
- Sume sync mode waits at most 30 seconds: short TTS vs long scripts
mode sync is a bounded wait, clamped to 0-30 seconds, not a promise the audio is ready. When to use it, and when to go async or webhook.
- Sume TTS 1.0 rejects model and model_id: use the router to pick
TTS 1.0 has no engine picker and returns 400 for model or model_id. The TTS router takes a required model from its catalog. Compare with ElevenLabs model tiers.
- Sume TTS emotion is a 64 character string: write a guide that fits
The emotion field in Sume TTS generation_config takes 1 to 64 characters. How to write a short, usable guide and test it against a neutral take.
- Sume TTS takes transcript or transcript_source, never both
A Sume TTS request accepts exactly one text input: literal transcript, or a transcript_source that points at a stored script. How to pick.
Written by Sume