Python exceptions for Sume errors: retry on the class, not the status

Map the Sume error envelope code to two exception classes, Retryable and Fatal, so one except clause drives retries. Covers 409, 429, 402 and 503.

5 min readSume
All posts

In Python, parse Sume's error envelope into two exception classes, Retryable and Fatal, keyed on error.code, and write your retry loop to catch only Retryable. The HTTP status alone cannot do this job: a 409 can mean wait (job_not_completed) or stop (idempotency_conflict), and a 429 can mean wait for the window (rate_limited) or wait for a queue slot (queue_full). Every Sume error comes in one shape, { "error": { "code", "message", "request_id", "details" } }, documented in Sume errors and credits, so the code is the stable thing to branch on.

A release tracker lists the OpenAI Sora API as ended on 2026-09-24 (MagicHour tracker, read 2026-10-06), so a lot of retry logic written for it is being rewritten right now. This is a good moment to replace status-based branching with something that survives a new error code, and to write the classification down in one table that a reviewer can read in a minute, instead of scattering it across catch blocks in five services.

Which codes are retryable?

From the error docs: rate_limited carries a retry-after header and ratelimit-* headers; queue_full is also a 429 but means your plan's queue is full; provider_capacity_exceeded (503) means retry later with the same key; provider_not_configured (503) means do not hammer it; job_not_completed (409) on a result or content fetch means keep polling. On the stop side are invalid_request (400), unauthorized (401), insufficient_credits (402), not_found and model_not_found (404), idempotency_conflict and job_not_cancelable (409), payload_too_large (413) and job_failed (409), where the next action is to read the failure from the job record.

What is the code?

class SumeError(Exception):
    def __init__(self, code, message, request_id=None, delay=None):
        super().__init__(f"{code}: {message}")
        self.code, self.request_id, self.delay = code, request_id, delay
class Retryable(SumeError): pass
class Fatal(SumeError): pass

WAIT = {"rate_limited", "queue_full", "provider_capacity_exceeded", "job_not_completed"}
def raise_for(status, headers, body):
    if status < 400: return
    err = body.get("error", {})
    code = err.get("code", "unknown")
    args = (code, err.get("message", ""), err.get("request_id"))
    if code in WAIT:
        after = headers.get("retry-after")
        raise Retryable(*args, delay=float(after) if after else None)
    raise Fatal(*args)

def attempt(responses):
    for status, headers, body in responses:
        try:
            return raise_for(status, headers, body) or "ok"
        except Retryable as e:
            print("retry in", e.delay or "backoff", e.code)
    raise Fatal("exhausted", "no success")
r429 = (429, {"retry-after": "7"}, {"error": {"code": "queue_full", "message": "x"}})
print(attempt([r429, (200, {}, {})]))
try: raise_for(409, {}, {"error": {"code": "idempotency_conflict"}})
except Fatal as e: print("stop", e.code)

Why two classes and not one per code?

Two rules the code encodes. First, a Retryable retry reuses the same Idempotency-Key, so a retry that lands after the first attempt quietly succeeded returns the original job instead of billing a second one. Second, Fatal is never retried by a loop; it goes to a human or to a branch that changes the request, like adding credits after a 402. insufficient_credits looks transient and is not: it stays until the balance changes, so a retry loop on it just burns your deadline.

Sample classification from Sume errors and credits (read 2026-10-06)
error.codeHTTPClassDelay source
rate_limited429Retryableretry-after header
queue_full429RetryableBackoff, or next poll hint
provider_capacity_exceeded503RetryableBackoff, same key
job_not_completed409Retryablenext_poll_after_seconds
idempotency_conflict409FatalFix the payload or the key
insufficient_credits402FatalAdd credits

How does the retry loop use them?

The loop that uses these classes stays short because the policy lives in the classification. Wrap the paid submit in for attempt in range(5), catch Retryable, sleep for e.delay when it is set and for a jittered backoff when it is not, and let Fatal propagate. Cap the total time as well as the count, because five waits of the maximum retry-after can outlast the request that triggered them, and a worker that holds a slot for minutes is its own capacity problem.

Keep the same Idempotency-Key across every attempt in that loop and across process restarts. Generate it from the business identity of the work, not from a fresh UUID per attempt, otherwise the safety net does nothing. For a 429 specifically, prefer the retry-after value over your own guess: the server knows when the window resets, and the ratelimit-reset header says the same thing in a different unit, so read one of them and not both.

What should an unknown code do?

Unknown codes should default to Fatal, not Retryable. A new code you have never seen is more likely a rule than a blip, and an unbounded retry on a paid POST is how a bug becomes a bill. Log request_id from the envelope with every raised error; it is what you hand to support when a job looks wrong, and it costs one field in the exception. If your service has a dashboard, count raised errors by class and by code, so a rise in Fatal with a single code tells you which deploy changed the request shape.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume