Free plan: 120 writes a minute but 6 accepted jobs, which hits first
Video batches on Sume hit queue_full long before the write rate limit. Per-plan arithmetic for 429 rate_limited versus queue_full, with a calculator.

For video batches on Sume, the accepted-job limit trips long before the write rate limit does. A Free workspace can send 120 writes a minute but can hold only 6 paid generation jobs, so the seventh simultaneous submit gets 429 queue_full, even though you are using 5 percent of your request budget. A Scale workspace has 1,200 writes a minute and 120 accepted jobs, with the same shape.
Two limits that share a status code
Both limits answer 429, which is why they get confused. rate_limited means request volume exceeded a window, and the response carries ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after. MDN defines retry-after as either a delay in seconds or an HTTP date, and Sume sends seconds.
queue_full means the workspace used all of its accepted generation capacity. It has nothing to do with requests per minute. It clears when a running or queued job finishes or is canceled, which for a 30-second video is minutes, not seconds.
The ratio per plan
Divide the write budget by the accepted capacity and you see how far apart the two limits sit. Reads have their own bucket at forty times the write number, so polling cannot starve your submits.
| Plan | Writes per minute | Reads per minute | Accepted jobs | Writes per accepted job |
|---|---|---|---|---|
| Free | 120 | 4,800 | 6 | 20 |
| Pro | 300 | 12,000 | 24 | 12.5 |
| Startup | 600 | 24,000 | 48 | 12.5 |
| Scale | 1,200 | 48,000 | 120 | 10 |
A calculator for your batch
Give it a plan and a count and it tells you how many waves you need, how much headroom a minute of writes gives you, and which limit is the binding one. It is plain Python, so you can run it before you spend anything.
A short worked case helps. Say you queue 40 Seedance 2.5 clips of 30 seconds on a Free plan. You can fire 40 POSTs in one minute without touching the 120-write budget, but only 6 are accepted at a time, so 34 of them return queue_full. If your code treats every 429 as a rate limit and sleeps for a fixed 60 seconds, you retry about once a minute and the slots sit idle between finishes. A pacer that submits when a job ends keeps the slots full.
PLANS = { # writes/min, accepted jobs (concurrency + queue)
"free": (120, 6), "pro": (300, 24), "startup": (600, 48), "scale": (1200, 120),
}
def plan_batch(plan: str, jobs: int, minutes_per_job: float = 3.0):
writes, accepted = PLANS[plan]
waves = -(-jobs // accepted) # ceiling division
binding = "queue_full" if accepted < writes else "rate_limited"
return {
"plan": plan,
"waves": waves,
"binding_limit": binding,
"submits_in_first_minute_allowed": min(jobs, writes),
"jobs_accepted_at_once": min(jobs, accepted),
"rough_wall_minutes": waves * minutes_per_job,
}
for p in PLANS:
print(plan_batch(p, 100))
Reading the output
For 100 clips the Free plan needs 17 waves and Pro needs 5. The binding limit is queue_full in every row, because accepted jobs are always far below writes per minute. The wall-clock figure is an input you choose, not a measurement, so replace minutes_per_job with what your own jobs take.
The practical advice is to size the job pacer around accepted capacity and let the rate limiter be a safety net. A pacer that sends one submit per slot that frees will never see rate_limited, and rarely see queue_full.
There is one more place the two limits meet: polling. Reads are a separate bucket at forty times the write number, so a status loop over 24 jobs every two seconds is about 720 reads a minute, well inside Pro's 12,000. You can poll generously and still keep your submit budget untouched, which is the reason the split exists.
What to do on each code
- 429 rate_limited: wait for retry-after, then resend with the same Idempotency-Key.
- 429 queue_full: wait for a job to finish, cancel queued jobs you do not need, then resend with the same key.
- Do not retry an unsafe submit without a key. A retry without one can bill twice.
- Check error.details.scope on a 429. It names the budget, read or write, that the request spent from.
Sources
Related posts
More in Developers
- Gemini 2.5 Flash Image shut down Oct 2: check old IDs on Sume
Google shut down gemini-2.5-flash-image on October 2, 2026 and points to Lite. On Sume an unknown model id returns 404 model_not_found; use a catalog id.
- Gemini Omni 1.1 Flash has no shutdown date yet: how to pin and watch
Google lists gemini-omni-1.1-flash with no shutdown date announced, while Veo 3.1 previews end October 22. Pin the id, and watch the deprecations page.
- Gemini Omni reference clips: 3 videos, 3 seconds each, VIDEO_REF_0
On Sume, Gemini Omni Flash reference-to-video takes up to 10 images and 3 clips of at most 3 s each, addressed as IMAGE_REF_0 and VIDEO_REF_0 in the prompt.
- Gemini Omni Flash edit on Sume: the 8-second reserve hint
A Gemini Omni edit has no duration input, so Sume reserves against a hint: 8 seconds by default, up to 30 if you send one. At 720p, that is $1.00 held.
Written by Sume