OpenAI 429 slow_down vs 503 server_is_overloaded, and Sume
OpenAI now splits 429 slow_down (traffic ramping too fast) from 503 server_is_overloaded. Here is how each maps to Sume's 429 scope and 503 retry rules.

OpenAI's changelog, read 2026-09-30, says traffic that increases too quickly can return a 429 with the code slow_down, while temporary model overload returns a 503 with server_is_overloaded. On Sume the two ideas live in different places: a 429 rate_limited names the budget in error.details.scope, and a provider_capacity_exceeded response means Sume's provider dispatch queue is full.
These are Sume's own error contracts. They do not translate OpenAI's codes, and Sume does not return slow_down or server_is_overloaded on this page's evidence.
What does OpenAI say to do?
The changelog says both responses may include Retry-After. When present, wait at least that long before retrying; when missing, use exponential backoff.
How does each case map to Sume?
Sume's rate limit docs give reads and writes separate budgets, so "a tight status-poll loop cannot 429 your own submits". A 429 says which budget it spent in error.details.scope, either read or write, and the response carries retry-after. For capacity, the errors page lists provider_capacity_exceeded.
| Situation | OpenAI | Sume |
|---|---|---|
| Requests too fast | 429 slow_down | 429, error.details.scope is read or write; wait retry-after |
| Provider capacity | 503 server_is_overloaded | provider_capacity_exceeded; retry later with the same idempotency key |
| Retry header | Retry-After may be present | retry-after is sent on 429 |
How should my retry code differ?
For a 429, read ratelimit-remaining and back off on retry-after instead of counting requests yourself. For provider_capacity_exceeded, retry later but keep the same Idempotency-Key, so a job that was in fact created is not created twice; see idempotency keys for AI video APIs.
async function withRetry(call, tries = 4) {
for (let i = 0; i < tries; i++) {
const res = await call();
if (res.status !== 429 && res.status !== 503) return res;
const wait = Number(res.headers.get("retry-after")) || 2 ** i;
await new Promise((r) => setTimeout(r, wait * 1000));
}
throw new Error("still rate limited or at capacity");
}Is every Sume 429 a request-rate problem?
No. queue_full also arrives as a 429 and means Sume cannot accept another paid job for the workspace until one finishes or is canceled; see queue full vs concurrency full. Request rate and generation concurrency are separate limits, and raising one does not raise the other.
Sources
Related posts
More in Developers
- OpenAI Agents SDK imageGenerationTool action edit and Sume
imageGenerationTool() takes action: 'generate' | 'edit' | 'auto' in openai-agents-js 0.18.0. To edit with Sume instead, wrap its image route as a function tool.
- OpenAI Agents SDK MCP require_approval for Sume write tools
Use require_approval with a tool_names list to gate Sume write tools by name, or connect with OAuth mcp:read so those tools are never visible.
- OpenAI Agents SDK MCP tool_input_guardrails for Sume spend
tool_input_guardrails on an Agents SDK MCP server can reject a call before it runs. For Sume paid tools, check idempotency_key and max_spend_usd.
- OpenAI Agents SDK MCPServerManager with a Sume server
Run Sume next to another MCP server in the OpenAI Agents Python SDK: connect Sume over streamable HTTP, check it with mcp_health, and read tools_list.
Written by Sume