Retry-After header: how long to wait after a 429 or 503
Retry-After tells a client how long to wait before retrying: a number of seconds or an HTTP date, sent with 429 or 503. How to read it and what to do.

Retry-After is an HTTP response header that tells a client how long to wait before its next request: either a number of seconds, as in Retry-After: 120, or an HTTP date. Servers can send it with 429 Too Many Requests and 503 Service Unavailable. Wait at least that long before you retry, and back off further if the retry fails too.
The definitions are quoted from RFC 9110 and RFC 6585, read 2026-09-28. The Sume examples come from its Authentication and Run webhooks docs, and from its current API code where marked.
Is Retry-After in seconds or milliseconds?
Seconds. RFC 9110 defines the value as an HTTP-date or delay-seconds, a non-negative decimal integer representing time in seconds; there is no millisecond form. Its two examples are Retry-After: Fri, 31 Dec 1999 23:59:59 GMT and Retry-After: 120, which means two minutes.
Handle both forms. This parser follows the same rules as the one in Sume's webhook sender in current code, where a date in the past means no wait:
// Wait in ms from a Retry-After value; null if absent or invalid.
function retryAfterMs(header: string | null, now = Date.now()): number | null {
const value = header?.trim();
if (!value) return null;
if (/^-?\d+$/.test(value)) {
const seconds = Number(value); // delay-seconds
return Number.isSafeInteger(seconds) && seconds >= 0 ? seconds * 1000 : null;
}
const date = Date.parse(value); // HTTP-date
return Number.isNaN(date) ? null : Math.max(0, date - now);
}How long does a 429 Too Many Requests last?
Until the server's rate limit lets you through again, and Retry-After is how the server tells you when that is. RFC 6585 defines 429 as too many requests in a given amount of time and says the response may include Retry-After; how requests are counted is up to the server.
On the Sume API, a 429 rate_limited means the key's per-minute read or write budget is spent, and its retry-after gives the seconds to wait. In current code the window is fixed per key and budget, retry-after is the seconds until it resets (at least 1), and the default window is 60 seconds, so a 429 rate_limited clears within a minute. Sume API errors and rate limits covers the budgets per plan and the ratelimit-* headers you can pace on.
Does every 429 and 503 include Retry-After?
No. RFC 9110 says a server may send it with a 503, and RFC 6585 says a 429 may include it, so code for its absence. Sume's 429 queue_full is a capacity limit rather than a rate limit, as video job concurrency and queueing explains, so it clears when jobs finish, not on a clock. In current code:
| Response | Wait hint | What to do |
|---|---|---|
429 rate_limited | retry-after header: seconds until the key's window resets | Wait for the window to reset, then resend. |
429 queue_full | retry_after_seconds: 30 in the body, with no header | Wait for running jobs to finish, or cancel queued jobs you don't need. A resend with the same Idempotency-Key gets the same queue_full back. |
503 deploy_draining | retry-after: 5 | Resend the same request after the wait. |
503 database_busy | retry-after: 1 | Resend after the wait. |
Does Sume honor Retry-After from my webhook endpoint?
For run webhooks, yes. If your endpoint answers a delivery with 429 or 503 and a Retry-After, Sume waits the longer of its own backoff and your value, capped at one hour. The docs give the schedule as min(max(30s × 2^(attempt−1) with jitter, Retry-After), 1h) over at most 10 attempts, and in current code the sender reads both the seconds form and the date form.
Job webhooks work differently: they retry on a fixed delay between attempts, 30 seconds by default.
What should I do when there is no Retry-After?
Back off exponentially with jitter and cap the attempts, so clients that failed together don't retry together. Sume's TypeScript SDK does this by default, with two retries, and honors retry-after when the response has one. Retryable HTTP status codes covers which errors deserve a retry at all, and when a POST is safe to resend.
Sources
Related posts
More in Developers
- Speech to text API in Java: transcribe audio with HttpClient
Call a speech to text API from Java with the JDK HttpClient and Jackson: post the audio URL, poll the job, then read the transcript and word times.
- Stateless MCP server: sessions, Mcp-Session-Id, and handles
A stateless MCP server keeps no session between requests. What Mcp-Session-Id does, what 2026-07-28 removed, and where state goes instead.
- Synthesia API documentation: the quickstart, step by step
Synthesia's API docs make a video in four steps: a Legacy (v2) key, POST /v2/videos, poll until complete, then download from a time-limited link.
- Text to image API: send a prompt, get image URLs back
A text-to-image API turns a prompt sent over HTTPS into generated images. How to call one: the request, the response, slow jobs, and the cost.
Written by Sume