429 vs 503: rate limit or server overload?

429 means you sent too many requests in a given time; 503 means the server can't handle requests right now. Both can carry Retry-After. What to do.

5 min readSume
All posts

A 429 Too Many Requests means you, the client, sent too many requests in a given amount of time, so slow down. A 503 Service Unavailable means the server is temporarily unable to handle the request, because of an overload or maintenance on its side. Both can carry a Retry-After header that says how long to wait before you try again.

The definitions are quoted from RFC 6585 and RFC 9110, read 2026-09-28. The Sume examples come from its Errors and rate limits and Generation admission docs, and from its current API code where marked.

What is the difference between 429 and 503?

Who is being limited. RFC 6585 defines 429 as a user sending too many requests in a given amount of time, which is rate limiting, and leaves it to the server how it identifies users and counts requests. RFC 9110 defines 503 as a server that can't handle the request because of a temporary overload or scheduled maintenance, which will likely be alleviated after some delay.

So a 429 is about your request rate, which you control, and a 503 about the server's state, which you don't. RFC 9110 also notes a server doesn't have to send 503 when overloaded: some simply refuse the connection.

What do 429 and 503 mean on the Sume API?

Each status carries more than one code, and they don't all clear the same way:

From Errors and rate limits, Generation admission, Authentication and current API code, read 2026-09-28.
Status and codeWhat it meansWhat to do
429 rate_limitedThe key's per-minute read or write budget is spent; error.details.scope says which.Wait retry-after seconds, then resend.
429 queue_fullThe workspace's generation concurrency plus queue capacity is full. It isn't a request-rate limit.Wait for jobs to finish, or cancel queued jobs you don't need.
503 provider_capacity_exceededSume's provider dispatch queue is full.Retry later.
503 provider_not_configuredProvider execution is unavailable in this runtime.Don't retry aggressively; check catalog and runtime status.
503 deploy_drainingCurrent code: the API replica is shutting down for a deploy.Resend the same request after retry-after, 5 seconds.
503 database_busyCurrent code: the API is briefly over capacity.Resend after retry-after, 1 second.

How long does a 503 last?

As long as the server needs. RFC 9110 only says the condition will likely be alleviated after some delay, and that the server may send Retry-After to suggest how long to wait. On Sume, the two 503s that carry retry-after in current code ask for seconds, not minutes, as the table shows.

Should I retry a 429 or a 503?

Usually yes, after waiting, as long as the request is safe to repeat, so send creates with an Idempotency-Key. The exception is an error body that says retryable: false: current code marks 503 provider_not_configured that way, and the docs say not to retry it aggressively. Reads and writes have separate budgets on Sume, so fast status polling can't 429 your own creates, and a 429 or 503 inside a poll loop is transient: the run keeps going. Sume API errors and rate limits lists the budgets per plan.

Watch the key on capacity errors. In current code, a create turned away with queue_full or provider_capacity_exceeded is stored as a failed job, and a resend with the same Idempotency-Key gets that error back, so don't count on a same-key retry succeeding once capacity frees up. Canceling queued jobs you no longer need frees capacity sooner; how to cancel a job shows the call.

Should my own API return 429 or 503?

Return 429 when one client sends too many requests in a given time, and 503 when the server itself can't handle requests right now. Send Retry-After with either when you know the wait, and per RFC 6585, include details explaining a 429.

If your endpoint receives Sume's run webhooks, both work as backpressure: answer a delivery with 429 or 503 and a Retry-After, and Sume waits the longer of its own backoff and your value, up to one hour. Retry-After header covers the header's two formats and how to read them.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume