Which limit stopped my agent: 402, queue_full or spend cap

Six different walls look alike from an agent loop. A diagnosis table that tells wallet, run spend cap, queue, request rate, scope and provider credits apart.

6 min readSume
All posts

When an agent stops getting results from Sume, one of six different limits is usually responsible, and they need opposite responses. A wallet that cannot cover the reserve, a per-run spend cap, a full queue, a request-rate limit, a key without the right scope and an exhausted provider account all show up as an error in the loop. Branch on the HTTP status and the code, never on the message, and the right next step falls out of the pair.

Six walls, six responses

The first three are about money and capacity. 402 insufficient_credits is raised at create, before anything runs: the wallet cannot reserve the estimate, and the response says next_action: add_funds. The run spend cap is different: it applies during the run, which then ends failed as format_run_failed, and the Formats error page tells you to compare usage.billable_amount_usd_micros with usage.generation_spend_cap_usd_micros before raising the brief. 429 queue_full is neither: it means the workspace's concurrency plus queue capacity is full, so you wait for a running job to finish.

The last three are about the request itself. 429 rate_limited is request volume, and error.details.scope says whether the read or the write bucket ran out. 403 insufficient_scope means the key was minted without the scope the call needs, and retrying in a loop is the costly mistake. provider_credits_exhausted is on Sume's side, not on your balance, with retryable: false.

read 2026-10-03
Where it showsCodeMeaningResponse
402insufficient_creditsWallet below the reserveHuman tops up in the dashboard
200 receipt, status failedformat_run_failedRun reached its spend capCompare billable with the cap
429queue_fullConcurrency plus queue fullWait or cancel queued jobs
429rate_limitedRequest volume, read or writeBack off with retry-after
403insufficient_scopeKey lacks formats:write or similarMint a new key
failed runprovider_credits_exhaustedProvider account, not yoursRetry later with a new key

A classifier for the loop

Put this decision in a function, not in the prompt. The snippet maps a status and code to the limit that fired, using only the fields the docs name. It runs as-is. In a real loop, the label would pick a branch: retry with backoff, stop and notify, or open a ticket with the request id from x-sume-request-id.

def which_limit(status, code, details=None):
    details = details or {}
    if status == 402:
        return "wallet balance (dashboard top-up)"
    if status == 429 and code == "queue_full":
        return "concurrency plus queue capacity (plan)"
    if status == 429 and code == "rate_limited":
        return f"request rate, {details.get('scope', 'unknown')} bucket"
    if code == "format_run_failed":
        return "run spend cap, if billable is near the cap"
    if code == "provider_credits_exhausted":
        return "Sume's provider account, not you"
    if status == 403 and code == "insufficient_scope":
        return "key scopes, fixed at mint time"
    return "not a limit: read the message and request id"

cases = [(402, "insufficient_credits"), (429, "queue_full"),
         (429, "rate_limited", {"scope": "write"}), (200, "format_run_failed"),
         (200, "provider_credits_exhausted"), (403, "insufficient_scope")]
for c in cases:
    print(c[1], "->", which_limit(*c))

Capacity numbers to have at hand

Concurrency is plan-only, and top-ups do not raise it: Free 1 processing job with a queue of 5, Pro 4 with 20, Startup 8 with 40, Scale 20 with 100. Submit rate limits are separate and count writes per minute: Free 120, Pro 300, Startup 600, Scale 1,200, with reads at 40 times that. When you hit queue_full, the generation_limits snapshot in the error details tells you how much room remains; the effective numbers on your dashboard Concurrency tab beat any static table.

Do not retry across classes

A retry policy that treats every 4xx the same wastes the budget. Transient classes (rate limit, queue full, provider capacity) deserve backoff and the same Idempotency-Key. Terminal classes (402, scope, spend cap) deserve a stop and a message to a person. Log the code and the request id on every failure so that the pattern, not the individual error, is what you look at.

Wiring the labels

Send the label to your logs and metrics as a dimension, so you can chart limits over time: a rising count of write-bucket rate limits says your pacing is off, a rising queue_full says you are submitting past the plan's accepted capacity, and a single 402 says a person has to act. Keep the request id with each so a support conversation is one paste. The same classifier can sit in front of an agent's tool result and replace a raw error body with a one-line instruction the model can follow without improvising.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume