Agent fans out 30 renders on a Pro plan: 24 accepted, 6 queue_full

Pro accepts 24 paid generation jobs at once (4 processing plus 20 queued). The 25th gets 429 queue_full. Retry policy for 402 and 429 in an agent loop.

5 min readSume
All posts

On a Sume Pro plan, an agent that submits 30 paid generation jobs at once gets 24 accepted and 6 rejected with 429 queue_full, because accepted capacity is 4 processing plus 20 queued. A 402 insufficient_credits is a different problem and should not be retried the same way.

The capacity table

The generation admission docs give the defaults. Queue capacity is max(3, concurrency_limit x 5), and accepted job capacity is the concurrency limit plus the queue limit. The docs also tell you to prefer the effective fields in the API response over the static table.

Generation concurrency and queue by plan, from the Generation admission page (read 2026-10-08)
PlanProcessingQueueAccepted
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

The fan-out

A fast planning model makes it easy to ask for many renders in one turn. Pro's 24 slots mean jobs 1 to 24 return as queued or processing, and jobs 25 to 30 fail with queue_full. The docs say full concurrency alone is not an error; it becomes one only when the queue is also full. The wave_size_hint in the response is max(1, floor(queue_capacity_remaining x 0.75)); for an empty Pro workspace that is floor(24 x 0.75) = 18, a submission-wave hint and not a limit.

Retry rules for an agent

The points that matter here, in the order you will hit them:

What an agent should do with each rejection, from the Generation admission page (read 2026-10-08)
CodeMeaningAgent action
429 queue_fullNo accepted capacity leftWait for jobs to finish or cancel queued ones; retry with the same idempotency key
429 rate_limitedRequest volume limitBack off using retry-after
402 insufficient_creditsCannot reserve the estimated costDo not retry; shrink the request or stop
409 idempotency_conflictKey reused with a different payloadUse a new key

Sizing a wave

The docs size new in-flight work as max(0, concurrency_limit - active - queued), limited to queue_capacity_remaining. Hosted MCP's run contract says the same in other words: fan out within headroom, wait on queue_full. Build that into the prompt or script of your agent, rather than letting the model discover the limit by failing.

Cost note

Admission is about capacity and balance, not model choice. Haiku 5.5 or GLM 5.3 Flash can plan the fan-out for cents, but each accepted job still reserves its estimated cost from your balance.

A second wave

After the first 24 are accepted, suppose 10 finish. The snapshot then shows 10 slots of capacity remaining, and the wave hint is floor(10 x 0.75) = 7. You can send 6 more jobs, which is exactly the 6 that were rejected, with the same idempotency keys. Because the keys are the same, a job that was in fact accepted before a network error is not duplicated.

Whatever the plan, read the live numbers. The docs say to prefer the effective concurrency_limit over the static table, because admin overrides can raise it and org workspaces have a floor of 10. Generation concurrency is plan-only: prepaid top-ups do not increase it. So if your agent regularly hits queue_full, the fix is a plan with more concurrency or a smaller wave, not a larger balance.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume