Which limit stopped my agent: 402, queue_full or spend cap
Six different walls look alike from an agent loop. A diagnosis table that tells wallet, run spend cap, queue, request rate, scope and provider credits apart.

When an agent stops getting results from Sume, one of six different limits is usually responsible, and they need opposite responses. A wallet that cannot cover the reserve, a per-run spend cap, a full queue, a request-rate limit, a key without the right scope and an exhausted provider account all show up as an error in the loop. Branch on the HTTP status and the code, never on the message, and the right next step falls out of the pair.
Six walls, six responses
The first three are about money and capacity. 402 insufficient_credits is raised at create, before anything runs: the wallet cannot reserve the estimate, and the response says next_action: add_funds. The run spend cap is different: it applies during the run, which then ends failed as format_run_failed, and the Formats error page tells you to compare usage.billable_amount_usd_micros with usage.generation_spend_cap_usd_micros before raising the brief. 429 queue_full is neither: it means the workspace's concurrency plus queue capacity is full, so you wait for a running job to finish.
The last three are about the request itself. 429 rate_limited is request volume, and error.details.scope says whether the read or the write bucket ran out. 403 insufficient_scope means the key was minted without the scope the call needs, and retrying in a loop is the costly mistake. provider_credits_exhausted is on Sume's side, not on your balance, with retryable: false.
| Where it shows | Code | Meaning | Response |
|---|---|---|---|
| 402 | insufficient_credits | Wallet below the reserve | Human tops up in the dashboard |
| 200 receipt, status failed | format_run_failed | Run reached its spend cap | Compare billable with the cap |
| 429 | queue_full | Concurrency plus queue full | Wait or cancel queued jobs |
| 429 | rate_limited | Request volume, read or write | Back off with retry-after |
| 403 | insufficient_scope | Key lacks formats:write or similar | Mint a new key |
| failed run | provider_credits_exhausted | Provider account, not yours | Retry later with a new key |
A classifier for the loop
Put this decision in a function, not in the prompt. The snippet maps a status and code to the limit that fired, using only the fields the docs name. It runs as-is. In a real loop, the label would pick a branch: retry with backoff, stop and notify, or open a ticket with the request id from x-sume-request-id.
def which_limit(status, code, details=None):
details = details or {}
if status == 402:
return "wallet balance (dashboard top-up)"
if status == 429 and code == "queue_full":
return "concurrency plus queue capacity (plan)"
if status == 429 and code == "rate_limited":
return f"request rate, {details.get('scope', 'unknown')} bucket"
if code == "format_run_failed":
return "run spend cap, if billable is near the cap"
if code == "provider_credits_exhausted":
return "Sume's provider account, not you"
if status == 403 and code == "insufficient_scope":
return "key scopes, fixed at mint time"
return "not a limit: read the message and request id"
cases = [(402, "insufficient_credits"), (429, "queue_full"),
(429, "rate_limited", {"scope": "write"}), (200, "format_run_failed"),
(200, "provider_credits_exhausted"), (403, "insufficient_scope")]
for c in cases:
print(c[1], "->", which_limit(*c))Capacity numbers to have at hand
Concurrency is plan-only, and top-ups do not raise it: Free 1 processing job with a queue of 5, Pro 4 with 20, Startup 8 with 40, Scale 20 with 100. Submit rate limits are separate and count writes per minute: Free 120, Pro 300, Startup 600, Scale 1,200, with reads at 40 times that. When you hit queue_full, the generation_limits snapshot in the error details tells you how much room remains; the effective numbers on your dashboard Concurrency tab beat any static table.
Do not retry across classes
A retry policy that treats every 4xx the same wastes the budget. Transient classes (rate limit, queue full, provider capacity) deserve backoff and the same Idempotency-Key. Terminal classes (402, scope, spend cap) deserve a stop and a message to a person. Log the code and the request id on every failure so that the pattern, not the individual error, is what you look at.
Wiring the labels
Send the label to your logs and metrics as a dimension, so you can chart limits over time: a rising count of write-bucket rate limits says your pacing is off, a rising queue_full says you are submitting past the plan's accepted capacity, and a single 402 says a person has to act. Keep the request id with each so a support conversation is one paste. The same classifier can sit in front of an agent's tool result and replace a raw error body with a one-line instruction the model can follow without improvising.
Sources
Related posts
More in Developers
- Which Sume API errors should page an engineer: route by category
Route Sume API failures by category: fix-the-input errors go to the caller, quota to finance, queue to a retry, and only internal or unexpected 5xx to on-call.
- Which Sume image models accept the quality parameter?
Only five Sume image rows list quality: ChatGPT Image 2, both ChatGPT Image 2.5 ids, Ideogram V3 and Ideogram 4.5. Every other row returns 400 if you send it.
- Windmill run_wait_result vs a Sume async submit: pick one per job
Windmill advises async mode and offers run_wait_result for short jobs. Sume mirrors that split: async submit plus polling for long work, sync only for short.
- Windmill sync returns 200 on error: check Sume's failed flag too
Windmill's sync webhook returns HTTP 200 with the error as JSON by default. Do not trust the status code alone for Sume jobs: read terminal, failed and sync.
Written by Sume