429 backoff with jitter in Python: OpenAI's advice, Sume's headers
OpenAI recommends exponential backoff with random jitter on 429. A Python status poller that honors Sume's retry-after first and falls back to backoff.

Wait the retry-after seconds when the 429 carries them, and use exponential backoff with random jitter when it does not. OpenAI's rate-limit guide describes exponential backoff as waiting briefly after a failure and increasing the delay after each retry, and says random jitter stops clients retrying at the same moment. Sume's docs say to use retry-after when present.
OpenAI facts are from its Rate limits guide (listed under Sources) and Sume facts from Errors and rate limits, read 2026-09-30.
What does each guide tell a client to do?
OpenAI says a 429 may include a Retry-After header giving seconds to wait. Sume says to back off on 429, use retry-after when present, and not retry unsafe submit requests without an Idempotency-Key.
| Rule | OpenAI | Sume |
|---|---|---|
| Wait signal | Retry-After header, may be present | retry-after when present |
| Fallback | Exponential backoff | Back off; exponential for polling |
| Avoid herding | Random jitter | SDK polls are jittered |
| Paid submit retry | Not covered here | Only with an Idempotency-Key |
How do I poll a Sume job with that policy?
Reads are cheap on Sume, but a 429 on a status read means the read failed, not the job. The job keeps running and billing, so retry the read instead of resubmitting.
import asyncio
import os
import random
import httpx
async def get_status(client: httpx.AsyncClient, job_id: str, attempts: int = 5):
for attempt in range(attempts):
r = await client.get(f"https://api.sume.com/v1/jobs/{job_id}/status")
if r.status_code != 429:
r.raise_for_status()
return r.json()
wait = float(r.headers.get("retry-after", 2**attempt))
await asyncio.sleep(wait + random.uniform(0, 1))
raise RuntimeError("still rate limited")
async def main():
headers = {"x-api-key": os.environ["SUME_API_KEY"]}
async with httpx.AsyncClient(headers=headers) as client:
print(await get_status(client, "job_123"))
asyncio.run(main())Does the Sume SDK already do this?
Yes for TypeScript: createSumeClient retries 408, 429, 5xx and transport failures twice by default with exponential backoff and jitter, honoring retry-after. A POST is retried only when it carries an Idempotency-Key. The run helpers also tolerate six consecutive transient read failures before giving up.
What are the limits of this answer?
Keep the loop bounded and log the x-sume-request-id of the last failure. A timeout in your poller does not cancel the job; store the job id and read it again later.
Sources
Related posts
More in Developers
- OpenAI Agents SDK MCPServerManager with a Sume server
Run Sume next to another MCP server in the OpenAI Agents Python SDK: connect Sume over streamable HTTP, check it with mcp_health, and read tools_list.
- OpenRouter video provider.options on Sume: rejected, not dropped
OpenRouter lists provider passthrough configuration. Sume v1 runs one backend per model, so a non-empty provider.options returns 400 unsupported_parameter.
- OpenRouter video seed on Sume: no v1 model accepts it
OpenRouter lists seed for deterministic video generation. On Sume no v1 video model accepts seed: each reports seed false and the field is rejected.
- OpenRouter video size parameter on Sume: 400 unsupported_parameter
On Sume every video model reports supported_sizes null, so a size parameter returns 400 unsupported_parameter. Send resolution and aspect_ratio instead.
Written by Sume