Claude Code 529 retry delay setting is not a Sume 429/503 retry
Claude Code 2.1.290 added a base delay setting for 529 retries. It does not govern Sume calls; follow retry-after and keep the Idempotency-Key.

CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS sets the base delay for Claude Code's own retries of an overloaded (529) request, and it has no effect on how a Sume call is retried. The Claude Code changelog for 2.1.290 (October 5, 2026) says it was added to set a longer base delay for the backoff when retrying an overloaded (529) request. Sume answers with 429 and 503, and tells you how long to wait.
Two different retry loops
The Claude Code line comes from the Claude Code changelog (read 2026-10-07). The Sume facts come from Errors and rate limits and Authentication.
A status-to-action function
The loop you write for Sume should read the response, not a client setting. This function maps a status and code to the next step.
def action(status, code, headers=None):
"""Map a Sume HTTP error to what the client should do next."""
headers = headers or {}
if status == 429 and code == "queue_full":
return "wait: a queued or processing job must finish or be canceled, then resend with the same Idempotency-Key"
if status in (429, 503) and code != "provider_not_configured":
wait = headers.get("retry-after", "backoff")
return f"retry after {wait} with the same Idempotency-Key"
table = {
400: "fix the request; do not retry",
401: "fix the credential (send exactly one of x-api-key or Authorization)",
402: "add funds or lower the cost; do not retry",
404: "check the id and the workspace; do not retry",
409: "read the job; do not resubmit",
413: "shrink the body; do not retry",
415: "send application/json; do not retry",
503: "do not retry aggressively; check runtime status",
}
return table.get(status, "log the request_id and stop")
if __name__ == "__main__":
print(action(429, "rate_limited", {"retry-after": "7"}))
print(action(429, "queue_full"))
print(action(503, "provider_capacity_exceeded"))
print(action(409, "job_not_completed"))What Sume tells you
The table lists the Sume signals that the code reads.
| Signal | Meaning | Do this |
|---|---|---|
429 with retry-after | Request budget used | Wait that long, then resend with the same key |
429 with queue_full | Concurrency and queue both full | Wait for a job to finish or cancel one; do not hammer |
503, provider_capacity_exceeded | Provider dispatch queue full | Retry later with the same key |
ratelimit-remaining | Budget left in the window | Slow down before it reaches 0 |
Keep the key
The docs say not to retry unsafe submit requests without an Idempotency-Key. With the key, a retry after a 429 or a lost reply returns the original job and does not bill twice. Reads and writes have separate budgets, so a status-poll loop cannot cause a 429 on your own submits.
Sources
Related posts
More in Developers
- Codex 0.160.0 MCP status for one server: confirm Sume with mcp_health
Codex 0.160.0 adds single-server MCP status discovery with thread connection reuse. Confirm the Sume session itself with mcp_health and tools_list.
- Codex 0.160.1 remote stdio MCP fix: does it affect Sume?
Codex 0.160.1 preserves SYSTEMROOT, TEMP and TMP for remote stdio MCP launches on Windows hosts. Sume's hosted MCP is HTTP, so no process is launched.
- communication.webhook_url 400: HTTPS, public host, 2048 chars
A communication.webhook_url that is not public HTTPS, is over 2048 characters, or points at localhost or a private network returns 400 invalid_request.
- How long do 100 AI video clips take? Concurrency slots by plan
100 jobs take ceil(100 / slots) waves: 100 on Free, 25 on Pro, 13 on Startup, 5 on Scale. Multiply by one job's time. Math and a runnable snippet.
Written by Sume