Python requests retry: backoff, Retry-After, and POST
Requests doesn't retry by default. Mount urllib3's Retry on a Session with backoff and status_forcelist, and retry POST only with an Idempotency-Key.

To retry with Python Requests, mount an HTTPAdapter(max_retries=Retry(...)) on a Session: urllib3's Retry sets how many retries, which status codes force one (status_forcelist), and how long to back off (backoff_factor), and it honors Retry-After by default. Requests does not retry failed connections on its own, and Retry leaves POST out of its default methods. Add POST only when every POST carries an Idempotency-Key, so a resent create returns the original run instead of charging twice.
Requests facts come from its Advanced Usage and Developer Interface pages, urllib3 facts from its urllib3.util reference, and Tenacity facts from its docs. Sume facts come from Create a run, Errors and spend, Errors and rate limits and Authentication. All were read on 2026-09-28. The calls are plain HTTPS; there is no Sume plugin involved.
How do I add retries to a requests Session?
Build one Retry, wrap it in an HTTPAdapter, and mount the adapter on the https:// prefix; every request through that session then follows the policy. This one retries 429, 500, 502, 503 and 504 with backoff and jitter, and includes POST only because every create sends an Idempotency-Key. Requests has no timeout unless you set one, so pass timeout= on every call as well.
import os
import requests
from requests.adapters import HTTPAdapter
from urllib3.util import Retry
retry = Retry(
total=4,
backoff_factor=1,
backoff_jitter=0.5,
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods={"GET", "HEAD", "POST"}, # POST only because every POST sends a key
raise_on_status=False, # return the last response instead of raising
)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=retry))
session.headers.update({"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"})
resp = session.post(
"https://api.sume.com/v1/formats/acme/weekly-promo/runs",
json={"input": {"week": "2026-W40"}},
headers={"Idempotency-Key": "weekly-promo-2026-W40"},
timeout=(3.05, 30),
)
resp.raise_for_status()
run = resp.json()["data"] # 202: new run. 200: replay, idempotency_hit is true.Which Retry settings matter, and what are the defaults?
With the default respect_retry_after_header=True, a 413, 429 or 503 that carries Retry-After triggers a retry and the wait follows the header, which Sume sends, in seconds, on a 429. Everything else waits on backoff_factor.
| Setting | Default | What it does |
|---|---|---|
total | 10 | Total retries allowed; takes precedence over the other counts |
allowed_methods | DELETE, GET, HEAD, OPTIONS, PUT, TRACE | Methods that are retried. POST is not in the default set |
status_forcelist | None | Status codes that force a retry for methods in allowed_methods |
backoff_factor | 0, no backoff | Sleeps factor × 2^(previous retries): 0.1 gives 0.0, 0.2, 0.4, 0.8 s … |
backoff_max | 120 | Longest single sleep, in seconds |
backoff_jitter | 0.0 | Adds a random 0 to n seconds to each sleep |
respect_retry_after_header | True | Honors Retry-After on 413, 429 and 503 |
raise_on_status | True | Raises instead of returning the last response when status retries run out |
Should I retry a POST with urllib3 Retry?
Only when the server can't run it twice. A connect error happens before the request is sent, but a read error happens after the request was sent to the server, and urllib3's docs warn that such a request may have side effects: a paid create might already be running. Requests' own retry example puts POST in allowed_methods; copy that only with an Idempotency-Key on every POST, since Sume's docs say not to retry unsafe submit requests without one. With a key, the same key and body get 200 and the original run, with no second charge; Idempotency keys for AI video APIs lists the other replay answers. Two cases shape the Retry setup:
- A read-timeout retry can reach Sume while the first request is still in flight and get
409 idempotency_key_in_use, retryable after about a second. Keep409out ofstatus_forcelistanyway: it also meansidempotency_conflict, a key reused with a different body, which must not be retried as is. - A timeout ends your wait, not the work. A create that timed out may already have started a run, and Sume's docs say abandoning a poll loop does not stop a run or its spend. With POST in
allowed_methods,Retryresends the create with the same key and body, and Sume returns that run instead of starting another.
Which errors should not be retried?
Most other 4xx. At create, a 4xx means nothing ran and nothing was charged, so fix the call. status_forcelist sees only the status, so it would also resend Sume's 502 attachment_fetch_failed, whose next_action is fix_input. When that matters, decide in your own code from the error envelope, which carries retryable and retry_after_seconds. Axios retry with an idempotency key shows that pattern in Node.
What about Tenacity?
Tenacity retries a Python function rather than an HTTP request. Its bare @retry retries forever without waiting, so always set a stop and a wait, such as @retry(stop=stop_after_attempt(5), wait=wait_random_exponential(multiplier=1, max=60)). It resends whatever the function sends, so the POST rule is the same: same key, same body.
Sources
Related posts
More in Integrations
- Rails webhooks: verify an HMAC signature in a controller
Read request.raw_post, skip CSRF for that action only, check OpenSSL::HMAC.hexdigest against each sume-v1 entry with secure_compare, then head 204.
- React Native API integration: call an AI video API safely
In React Native, call your own backend with fetch and let it hold the API key. React Native's docs warn that anything in the app bundle is readable.
- Remotion with Claude Code: add AI video, voice and music
Set up Remotion's Agent Skills for Claude Code, then load generated clips, narration and music as files, with TTS word timings as captions.
- Replicate MCP server: remote and local setup for Claude
Replicate's MCP server is hosted at mcp.replicate.com or runs locally with npx replicate-mcp. Both use a Replicate API token. Setup per client.
Written by Sume