Which Sume submit errors to retry: a Python status triage function

Retry 429 and 503 with the same Idempotency-Key, fix 400, 413 and 415, stop on 402, poll on 409 job_not_completed. A short Python function and its table.

4 min readSume
All posts

Retry a Sume submit only when the status says waiting will help: 429 rate_limited, 429 queue_full and 503 provider_capacity_exceeded all retry, and they retry with the same Idempotency-Key. Fix 400, 413 and 415 first, stop on 402, and treat 409 as a state problem to read, not a request to repeat.

A small function that maps status and error code to one of a handful of actions keeps that logic in one place instead of scattered through every call site.

The mapping the docs give

The codes below come from the Sume pages on errors and on generation admission (read 2026-10-03).

Sume submit errors and the client action (read 2026-10-03)
Status and codeMeaningAction
400 invalid_requestBody, query, path or headers are invalidFix the request, then retry
401 unauthorizedKey missing or invalidFix authentication
402 insufficient_creditsBalance cannot reserve the estimated costAdd funds or lower the cost; do not loop
404 not_foundNot in the current workspaceCheck the id and the key's workspace
409 idempotency_conflictSame key reused for a different payloadUse a new key only for a new payload
413 payload_too_largeBody over the API limitShrink the body
415 unsupported_media_typeBody was not application/jsonSend JSON
429 rate_limitedRequest volume over the abuse limitWait retry-after, retry with the same key
429 queue_fullNo accepted-job capacity leftWait for jobs to finish or cancel queued ones, retry with the same key
503 provider_capacity_exceededDispatch queue fullRetry later with the same key

The function

It returns a short instruction string. In a real client you would return an enum and attach a sleep for the retry cases, but the branching is the point. The sample runs as written and prints four lines.

def triage(status: int, code: str = "") -> str:
    if status in (400, 413, 415):
        return "fix the request, do not retry"
    if status in (401, 403):
        return "fix the key or its scope"
    if status == 402:
        return "stop: add funds or lower the cost"
    if status == 404:
        return "check the id, workspace and key owner"
    if status == 409:
        if code == "job_not_completed":
            return "poll status_url, then read the result"
        if code == "idempotency_conflict":
            return "new payload needs a new Idempotency-Key"
        return "read the job state, do not resubmit"
    if status == 429:
        if code == "queue_full":
            return "wait for jobs to finish, retry with the same key"
        return "sleep retry-after, retry with the same key"
    if status == 503:
        return "retry later with the same key"
    return "log request_id and escalate"

for case in [(429, "queue_full"), (409, "idempotency_conflict"), (402, ""), (503, "provider_capacity_exceeded")]:
    print(case, "->", triage(*case))

Two details that save money

Reuse the same Idempotency-Key on every retry of the same request. A retry that arrives after the original was accepted returns the original job, so a timeout followed by a retry does not bill twice. A different key for the same intent is a second job.

Do not retry 402 in a loop. A 402 means the reservation failed before provider work started, so repeating the call only burns request budget. And keep 409 out of the retry path: a 409 job_not_completed on a result read means poll status_url, and the job itself is fine.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume