429 queue_full vs rate_limited: what the generation API means

On Sume, 429 queue_full means no accepted generation capacity is left, rate_limited means request volume, and full concurrency leaves the job queued.

4 min readSume
All posts

Both are 429, but they mean different things. On Sume, queue_full means the workspace has no remaining accepted generation capacity, so wait for jobs to finish or cancel queued ones. rate_limited means request volume passed an abuse-protection limit, so back off. A full concurrency limit is neither: the job simply stays queued.

How do the codes differ?

Generation submit responses, from the Sume generation admission docs, read 2026-09-29.
Status and codeWhy it happensWhat to do
429 queue_fullThe workspace has no remaining accepted generation capacityWait for jobs to finish or cancel queued jobs, then retry with the same idempotency key
429 rate_limitedAPI request volume exceeded an abuse-protection limitBack off using retry-after when present
402 insufficient_creditsSume cannot reserve the estimated cost from the balanceReduce the request or add spend capacity

Is a full concurrency limit an error?

No. Concurrency is a dispatch limit, not a submit limit. If your workspace is at its generation concurrency limit, Sume can still accept more jobs as queued while queue capacity remains, and workers move them to processing later. Concurrency being full becomes a submit error only when the queue is also full. Do not treat queued as failure: store the job_id, poll status with backoff, and fetch the result when it is ready.

How big is the queue?

Queue capacity defaults to max(3, concurrency_limit × 5). Accepted job capacity is the concurrency limit plus the queue, so on the default plan table a Free workspace holds 6 jobs and a Pro workspace 24. The docs say to prefer the effective generation_limits fields over the static table.

Default processing and queue limits, from the Sume docs, read 2026-09-29.
PlanProcessingQueueAccepted
Free156
Pro42024
Startup84048
Scale20100120

How do I retry safely?

Reuse the same idempotency key on an exact retry, which is what the docs prescribe for queue_full. Reusing a key for a different payload returns 409 idempotency_conflict. Read and status endpoints have their own limits; treat those as polling backpressure, not concurrency.

Where do I read my real limits?

Generation submit responses include a generation_limits snapshot when Sume can compute it, and the docs advise checking it and stopping new work when queue capacity is low. The dashboard Concurrency tab is the source of truth for the configured processing cap. Concurrency is plan-only: prepaid top-ups do not raise it, and org workspaces have a floor of 10.

Should I cancel queued jobs?

Only if you no longer want them. For queue_full, the docs give two ways out: wait for jobs to finish, or cancel queued jobs, then retry with the same idempotency key. Do not resubmit a paid request just because your local worker timed out; store status_url, result_url, events_url and cancel_url when present so you can act on the job you already have.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume