429 queue_full vs rate_limited: what the generation API means
On Sume, 429 queue_full means no accepted generation capacity is left, rate_limited means request volume, and full concurrency leaves the job queued.

Both are 429, but they mean different things. On Sume, queue_full means the workspace has no remaining accepted generation capacity, so wait for jobs to finish or cancel queued ones. rate_limited means request volume passed an abuse-protection limit, so back off. A full concurrency limit is neither: the job simply stays queued.
How do the codes differ?
| Status and code | Why it happens | What to do |
|---|---|---|
429 queue_full | The workspace has no remaining accepted generation capacity | Wait for jobs to finish or cancel queued jobs, then retry with the same idempotency key |
429 rate_limited | API request volume exceeded an abuse-protection limit | Back off using retry-after when present |
402 insufficient_credits | Sume cannot reserve the estimated cost from the balance | Reduce the request or add spend capacity |
Is a full concurrency limit an error?
No. Concurrency is a dispatch limit, not a submit limit. If your workspace is at its generation concurrency limit, Sume can still accept more jobs as queued while queue capacity remains, and workers move them to processing later. Concurrency being full becomes a submit error only when the queue is also full. Do not treat queued as failure: store the job_id, poll status with backoff, and fetch the result when it is ready.
How big is the queue?
Queue capacity defaults to max(3, concurrency_limit × 5). Accepted job capacity is the concurrency limit plus the queue, so on the default plan table a Free workspace holds 6 jobs and a Pro workspace 24. The docs say to prefer the effective generation_limits fields over the static table.
| Plan | Processing | Queue | Accepted |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
How do I retry safely?
Reuse the same idempotency key on an exact retry, which is what the docs prescribe for queue_full. Reusing a key for a different payload returns 409 idempotency_conflict. Read and status endpoints have their own limits; treat those as polling backpressure, not concurrency.
Where do I read my real limits?
Generation submit responses include a generation_limits snapshot when Sume can compute it, and the docs advise checking it and stopping new work when queue capacity is low. The dashboard Concurrency tab is the source of truth for the configured processing cap. Concurrency is plan-only: prepaid top-ups do not raise it, and org workspaces have a floor of 10.
Should I cancel queued jobs?
Only if you no longer want them. For queue_full, the docs give two ways out: wait for jobs to finish, or cancel queued jobs, then retry with the same idempotency key. Do not resubmit a paid request just because your local worker timed out; store status_url, result_url, events_url and cancel_url when present so you can act on the job you already have.
Sources
Related posts
More in Developers
- Real time speech to text API: what a file-based API can do
Sume's speech to text API isn't real time: it transcribes recordings of up to 10 minutes at a URL. Chunked recordings give near-live transcripts.
- Extract on-screen text from a reference video with an API
Reference ingest reads on-screen text at source resolution, returns lines with boxes, spans and confidence, and flags low-confidence lines instead of guessing.
- Reference video shot breakdown API: cuts, keyframes and timings
Break a reference video into shots with an API: POST /v1/reference-ingest returns frame-exact cuts, keyframes and a labeled strip for one clip, unbilled.
- Check whether a reference video is silent before adding music
Reference ingest reports audio.silent at a -60 LUFS gate, speech presence and beats, so you know whether to keep, replace or add a soundtrack before a remix.
Written by Sume