Retryable HTTP status codes: which errors to retry

Retry network errors, 408, 429, and 5xx with backoff; skip most other 4xx. Retry a POST only with an idempotency key, and read the API's retry flag.

5 min readSume
All posts

Retry a request when the failure is likely temporary: no response at all (a timeout or a dropped connection), 408 Request Timeout, 429 Too Many Requests, and 5xx server errors such as 500, 502, 503, and 504. Don't retry most other 4xx errors, which say the request itself is wrong: the same request gets the same answer. Retry a POST only when it carries an idempotency key, so a resend can't do the work twice.

Status definitions are quoted from RFC 9110 and RFC 6585, read 2026-09-28. The Sume examples come from its Errors and spend, Waiting for runs and jobs and Create a run docs, and from its current API code where marked.

Which HTTP status codes should be retried?

The ones where the same request can succeed a moment later. RFC 9110's classes point the way: a 4xx means the client seems to have erred, and a 5xx means the server erred or can't perform the request. Sume's TypeScript SDK draws the line in the same place: it retries 408, 429, 5xx, and transport failures.

From RFC 9110, RFC 6585 and Sume's Waiting for runs and jobs docs, read 2026-09-28.
StatusRetry?Why
No response (timeout, reset)Yes, if the request is safe to repeatYou can't tell whether the server applied it.
408 Request TimeoutYesThe server didn't receive the whole request in time; RFC 9110 says the client may repeat it.
429 Too Many RequestsYes, after Retry-AfterRate limiting; the response may say how long to wait.
500 Internal Server ErrorYes, a few times, if the request is safe to repeatAn unexpected condition on the server.
502 Bad Gateway, 504 Gateway TimeoutYes, if the request is safe to repeatA gateway got a bad or late answer from upstream, which may still have acted.
503 Service UnavailableYes, after Retry-AfterA temporary overload or maintenance, likely alleviated after some delay.
400, 401, 403, 404, 405, 413, 415, 422NoThe client erred; fix the request first.
409 ConflictDepends on the APISome conflicts clear in a moment; others never do.

Is it safe to retry a POST?

Only if a resend can't do the work twice: RFC 9110 says a client should not automatically retry a non-idempotent request, such as a POST, unless it knows the request is actually idempotent. An idempotency key is how an API makes a POST safe to resend; on Sume, the same Idempotency-Key with the same body returns the original job or run instead of a second one. Is POST idempotent? covers the method rules, and idempotency keys for AI video APIs covers key design.

When does the status code alone mislead?

When the error body says otherwise. RFC 9110 asks a server to explain a 4xx or 5xx in the response, including whether the condition is temporary or permanent. Sume's error body does that with retryable and retry_after_seconds, which say whether resending the same request can succeed and how long to wait first. Four cases where the flag and a status-code rule disagree; in the last one, trust the job's status over the flag:

  • 502 attachment_fetch_failed: Sume couldn't fetch your attachment URL. Despite the 5xx it is your input, with next_action: fix_input, so make the URL publicly reachable instead of retrying.
  • 503 provider_not_configured: provider execution is unavailable in that runtime. The docs say not to retry aggressively, and current code marks it retryable: false.
  • 409 idempotency_key_in_use: another request with the same key is in flight. It is retryable: true; wait about a second and resend.
  • 409 job_not_completed from GET /v1/jobs/{id}/result: in current code it is marked retryable even when the job failed or was canceled, so poll the job's status until it is terminal. 409 Conflict covers the other conflicts.

Should you retry 500 errors?

Yes, a limited number of times with backoff, as long as the request is safe to repeat. In current code Sume marks an unexpected 500 retryable: true with next_action: contact_support: retry, and if it keeps failing, quote the request_id to support.

While you poll a Sume run, the docs call a 429 or 503 transient as well: abandoning the poll loop doesn't stop the run or its spend, so back off and poll again.

How long should I wait between retries?

Use Retry-After when the response has one, and exponential backoff with jitter when it doesn't. Sume's SDK does both by default, with two retries. Retry-After header explains the header's two formats and how long a 429 lasts.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume