429 vs 503: rate limit or server overload?
429 means you sent too many requests in a given time; 503 means the server can't handle requests right now. Both can carry Retry-After. What to do.

A 429 Too Many Requests means you, the client, sent too many requests in a given amount of time, so slow down. A 503 Service Unavailable means the server is temporarily unable to handle the request, because of an overload or maintenance on its side. Both can carry a Retry-After header that says how long to wait before you try again.
The definitions are quoted from RFC 6585 and RFC 9110, read 2026-09-28. The Sume examples come from its Errors and rate limits and Generation admission docs, and from its current API code where marked.
What is the difference between 429 and 503?
Who is being limited. RFC 6585 defines 429 as a user sending too many requests in a given amount of time, which is rate limiting, and leaves it to the server how it identifies users and counts requests. RFC 9110 defines 503 as a server that can't handle the request because of a temporary overload or scheduled maintenance, which will likely be alleviated after some delay.
So a 429 is about your request rate, which you control, and a 503 about the server's state, which you don't. RFC 9110 also notes a server doesn't have to send 503 when overloaded: some simply refuse the connection.
What do 429 and 503 mean on the Sume API?
Each status carries more than one code, and they don't all clear the same way:
| Status and code | What it means | What to do |
|---|---|---|
429 rate_limited | The key's per-minute read or write budget is spent; error.details.scope says which. | Wait retry-after seconds, then resend. |
429 queue_full | The workspace's generation concurrency plus queue capacity is full. It isn't a request-rate limit. | Wait for jobs to finish, or cancel queued jobs you don't need. |
503 provider_capacity_exceeded | Sume's provider dispatch queue is full. | Retry later. |
503 provider_not_configured | Provider execution is unavailable in this runtime. | Don't retry aggressively; check catalog and runtime status. |
503 deploy_draining | Current code: the API replica is shutting down for a deploy. | Resend the same request after retry-after, 5 seconds. |
503 database_busy | Current code: the API is briefly over capacity. | Resend after retry-after, 1 second. |
How long does a 503 last?
As long as the server needs. RFC 9110 only says the condition will likely be alleviated after some delay, and that the server may send Retry-After to suggest how long to wait. On Sume, the two 503s that carry retry-after in current code ask for seconds, not minutes, as the table shows.
Should I retry a 429 or a 503?
Usually yes, after waiting, as long as the request is safe to repeat, so send creates with an Idempotency-Key. The exception is an error body that says retryable: false: current code marks 503 provider_not_configured that way, and the docs say not to retry it aggressively. Reads and writes have separate budgets on Sume, so fast status polling can't 429 your own creates, and a 429 or 503 inside a poll loop is transient: the run keeps going. Sume API errors and rate limits lists the budgets per plan.
Watch the key on capacity errors. In current code, a create turned away with queue_full or provider_capacity_exceeded is stored as a failed job, and a resend with the same Idempotency-Key gets that error back, so don't count on a same-key retry succeeding once capacity frees up. Canceling queued jobs you no longer need frees capacity sooner; how to cancel a job shows the call.
Should my own API return 429 or 503?
Return 429 when one client sends too many requests in a given time, and 503 when the server itself can't handle requests right now. Send Retry-After with either when you know the wait, and per RFC 6585, include details explaining a 429.
If your endpoint receives Sume's run webhooks, both work as backpressure: answer a delivery with 429 or 503 and a Retry-After, and Sume waits the longer of its own backoff and your value, up to one hour. Retry-After header covers the header's two formats and how to read them.
Sources
Related posts
More in Developers
- 504 Gateway Timeout from an API: did my request go through?
A 504 from an API means a proxy stopped waiting for the server. Your request may still be running, so check before you resend. How to avoid 504s.
- AI model aggregator: what it is and when to go direct
An AI model aggregator sells many labs' models behind one API key, one request shape and one bill. What you gain, what you give up, and what to check.
- AI model router: what it does and what you give up
An AI model router picks which model serves each request, so your code sends one id. How that differs from pinning a model, and the control you trade.
- Why an API returns 404 Not Found, and how to fix it
An API returns 404 when no route matches your path or method, or when the id doesn't exist for your credentials. How to tell them apart and fix each.
Written by Sume