Python httpx 429 handler for Sume: retry-after, then ratelimit-reset

A small async Python wrapper for the Sume API that waits on retry-after, falls back to ratelimit-reset, and never retries a POST that lacks an Idempotency-Key.

5 min readSume
All posts

When the Sume API answers 429, wait for the number of seconds in retry-after; if the header is missing, ratelimit-reset tells you how long until the window resets. Retry a GET freely. Retry a POST only if it carries an Idempotency-Key, because without one a retry can start and bill a second job. The wrapper below does exactly this with httpx.

The rules come from the Errors and rate limits page, which says to back off on 429, use retry-after when it is present, and not retry unsafe submit requests without a key. The Authentication page defines the four headers.

The headers

Public API responses can include these headers. The meanings are quoted from the Authentication page, read 2026-10-09.

Rate-limit headers on Sume responses, as of 2026-10-09 (Authentication).
HeaderMeaningWhen you use it
ratelimit-limitRequests permitted in the current windowLogging, dashboards
ratelimit-remainingRequests left in the current windowSlow down before you reach 0
ratelimit-resetSeconds until the window resetsFallback wait when retry-after is absent
retry-afterSeconds to wait, sent on a 429First choice on a 429

The wrapper

The wrapper takes any method and path, and sleeps before it tries again. Python code that awaits must run inside asyncio.run, so the entry point is main(). The final line is a read of GET /v1/balance, which is a documented endpoint.

import asyncio
import os

import httpx

BASE = "https://api.sume.com"


async def call(client, method, path, tries=5, **kw):
    headers = kw.setdefault("headers", {})
    safe_to_retry = method == "GET" or "Idempotency-Key" in headers
    for attempt in range(tries):
        r = await client.request(method, BASE + path, **kw)
        if r.status_code != 429 or not safe_to_retry:
            return r
        wait = r.headers.get("retry-after") or r.headers.get("ratelimit-reset")
        await asyncio.sleep(float(wait) if wait else 2**attempt)
    return r


async def main():
    key = os.environ["SUME_API_KEY"]
    async with httpx.AsyncClient(headers={"Authorization": f"Bearer {key}"}) as client:
        r = await call(client, "GET", "/v1/balance")
        print(r.status_code, r.headers.get("ratelimit-remaining"))


asyncio.run(main())

A 429 is not always the request rate

Two different limits answer with 429. rate_limited means you went over the per-minute budget, and error.details.scope names the bucket, read or write. queue_full means the workspace has no accepted generation capacity left. Waiting a second does not change that: wait for a job to finish or cancel queued jobs, then retry with the same Idempotency-Key. The wrapper handles both only in the sense that it waits for the header value. Add a check on error.code if you want to stop early on queue_full.

Jitter matters when many workers share one key. If twenty processes all sleep for the same retry-after value, they wake together and hit the same budget again. Add a random fraction of a second to the sleep, or stagger the start of each worker. The SDK's own client does something similar for TypeScript users: it retries 408, 429, and 5xx with exponential backoff and jitter, honors retry-after, and retries a POST only when it carries an Idempotency-Key.

Cap the number of tries, as the snippet does. A loop that never ends hides a real problem such as a runaway poller. If you see read-bucket 429s, lengthen the poll interval and add jitter before you ask for a higher plan.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume