429 rate_limit_unavailable: not your traffic, retry in Python
Sume can answer 429 rate_limit_unavailable when its limiter is degraded. It is not rate_limited and needs a full-window wait. Python code tells them apart.

Most 429s from the Sume API mean you spent a budget. rate_limited says so: the body's details carry limit, remaining, a scope of read or write, window_seconds and retry_after_seconds. There is a second 429 that looks the same and is not your fault. When the shared store that counts requests is unavailable, the API can answer 429 rate_limit_unavailable instead.
Whether that happens is a deployment policy. The API's rate limiter has a fail-open mode and a fail-closed mode, and the production default with Redis configured is fail-closed. In that mode every request that cannot be counted is refused, so a store outage shows up to callers as a burst of 429s that no amount of slowing down fixes.
How to tell the two apart
| Field | rate_limited | rate_limit_unavailable |
|---|---|---|
| Meaning | A key or IP spent its read or write budget | The limiter store is degraded and the policy fails closed |
details.rate_limit_degraded | absent | true |
| Error category and stage | rate_limit, queue | rate_limit, runtime |
| Retry hint | Seconds until the window resets | A full window length, at least 1 second |
| Does a retry help | Yes, after the wait | Yes, once the store recovers |
A retry loop that names the cause
The function below uses only the standard library. It retries only on 429, prefers the retry-after header and falls back to details.retry_after_seconds and then an exponential delay. It prints which kind of 429 it hit, which is the part worth putting in your logs. The last line reads a placeholder path, so replace it with a real read such as /v1/jobs/{id}/status.
import json, os, time, urllib.error, urllib.request
def get(url: str, tries: int = 4) -> dict:
for attempt in range(tries):
req = urllib.request.Request(url, headers={"x-api-key": os.environ["SUME_API_KEY"]})
try:
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)
except urllib.error.HTTPError as e:
if e.code != 429 or attempt == tries - 1:
raise
err = json.load(e).get("error", {})
d = err.get("details") or {}
if err.get("code") == "rate_limit_unavailable":
kind = "limiter degraded" # not your traffic
else:
kind = f"{d.get('scope', '?')} budget spent"
wait = float(e.headers.get("retry-after") or d.get("retry_after_seconds") or 2**attempt)
print(f"429 {err.get('code')}: {kind}; sleeping {wait:g}s")
time.sleep(wait)
print(get("https://api.sume.com/v1/jobs/deg"))What to do with the distinction
- For
rate_limited, readdetails.scope. Aread429 means your polling is too tight. Awrite429 means you are submitting too fast. The two budgets are separate, and reads default to 40 times the write budget. - For
rate_limit_unavailable, do not shrink your batch or open a spend alert. Wait the hinted window and retry. If it repeats for several windows, include thex-sume-request-idresponse header when you contact support. - Do not treat either as generation concurrency. A full queue is
queue_full, a different 429 that has its own retry rule. - Resubmitting a paid call is safe only with the same
Idempotency-Key, so keep one on every submit that your loop might retry.
The budgets themselves, by plan, are in the authentication guide. Polling advice, including when to stop, is in the jobs guide.
Sources
Related posts
More in Developers
- Read a completed Sume /v1/videos poll response, field by field
What id, generation_id, polling_url, status, unsigned_urls and usage.cost mean on a finished Sume video job, and which to store.
- Read capabilities from the Video Router models list before you pin
GET /v1/video-router/models returns capabilities per model. Why the docs say to read them instead of assuming one envelope, and what differs by model.
- Read job events for a stuck narration take: a snapshot, not a stream
GET /v1/jobs/:id/events lists job.created, queued, started, generation.submitted and the terminal event. A pull snapshot for debugging a TTS or music take.
- Debug a slow Sume job with GET /v1/jobs/:id/events
A slow Sume job is queued, running, or waiting on your webhook. The events timeline separates them: job.queued, job.started, terminal, webhook.delivery.
Written by Sume