Retry or fix? A status-code policy for Sume API clients in Python
Retry 429 and 503 with the same idempotency key; never retry 400, 401, 402 or 404. A small Python policy function, tested, with the Sume docs rules behind it.

Retry only what can change on its own: 429, 503, and a 409 idempotency_key_in_use after about a second. Fix, do not retry, anything that is your request or your account: 400, 401, 402, 404. And always retry a paid submit with the same Idempotency-Key, because without one the docs say not to retry unsafe submits at all.
That is a policy small enough to fit in one function, which is the point: a single place decides, so no call site invents its own loop.
The policy table
Rows come from the errors page and the idempotency rules. The "Wait" column is my choice of default, not a Sume value; Sume says to use retry-after when present.
| Status or code | Action | Wait |
|---|---|---|
429 rate_limited | Retry, same key | retry-after, else backoff |
429 queue_full | Retry, same key, after capacity opens | retry-after, else a longer backoff |
503 capacity or runtime | Retry, same key; do not hammer | Backoff with jitter |
409 idempotency_key_in_use | Retry, same key | About 1 s |
409 idempotency_conflict | Stop: body changed under a used key | None |
400, 401, 402, 404 | Fix the request or account | None |
Why not retry everything
A 402 insufficient_credits will not improve with time, so a retry loop only adds noise. A 400 unsupported_parameter is the API telling you a field is not supported by that model; Sume rejects it instead of dropping it silently. A 401 with both Authorization and X-API-Key set is solved by sending one.
Retries that do make sense need a bound. Cap attempts, add jitter so a fleet does not wake in step, and give up with a message that carries the request id.
The function
It returns the number of seconds to wait, or None for "do not retry". Wire it to your HTTP layer of choice.
import random
RETRY_SAME_KEY = {429, 503}
NO_RETRY = {400, 401, 402, 404}
def wait_seconds(status, code, retry_after, attempt):
if status in NO_RETRY or attempt >= 6:
return None
if status == 409:
return 1.0 if code == "idempotency_key_in_use" else None
if status in RETRY_SAME_KEY:
if retry_after is not None:
return float(retry_after)
return min(60.0, 2.0 ** attempt) * random.uniform(0.5, 1.0)
return None
print(wait_seconds(402, "insufficient_credits", None, 0))
print(wait_seconds(409, "idempotency_conflict", None, 0))
print(wait_seconds(409, "idempotency_key_in_use", None, 0))
print(wait_seconds(429, "queue_full", "12", 1))
print(0.5 <= wait_seconds(503, None, None, 2) <= 4.0)
Test it
The five print lines above are a test you can keep: None, None, 1.0, 12.0 and True. Add a case for each row of the table, and a final one for attempt 6, which must stop.
Where the policy lives
Put the function in your HTTP wrapper, not in each caller. The wrapper creates the idempotency key once per logical request, passes it on every attempt, and records the attempt count and the final status.
Keep the policy honest by reading the error body as well as the status. Sume errors carry a code, a request id and, for jobs, retryability and next-action fields, so a 503 that names a worker timeout can tell you to poll the status instead of resubmitting. When the body and your table disagree, trust the body.
Add a metric for each outcome: retried and succeeded, retried and gave up, and fixed (no retry). A rising count of the third kind means your own validation is behind the API, which is cheaper to fix than a growing retry budget.
- One key per logical request, not per attempt.
- Respect
retry-afterbefore your own backoff. - Cap total elapsed time, not just attempts.
Sources
Related posts
More in Developers
- sume/auto is missing from GET /v1/images/models: hardcode it
Sume Image Router Auto never appears in the model list. Discovery code that only reads GET /v1/images/models will never offer sume/auto. How to handle it.
- SUME_API_AUTH_MODE: make the Sume CLI send Bearer instead of x-api-key
The Sume CLI sends x-api-key by default. Set SUME_API_AUTH_MODE=bearer when your client or network layer expects Authorization: Bearer. Never send both.
- No progress percent in Sume job events: what to log and show instead
Sume's job events are a public timeline of eight named events, not a progress feed. What each means, what to log, and what to show a user while a video renders.
- A Sume job looks stuck: wait, cancel or poll the events?
Read status, then events. Queued and processing mean wait, cancel only works before generation starts, and a client timeout never cancels the job.
Written by Sume