Which statuses to retry when reading back a Sume job (Python)

After a Sume submit returns a job id, retry reads on 408, 425, 429, 500, 502, 503, 504 and 520 to 525. A tested Python classifier and the 409 gotcha.

4 min readSume
All posts

Once a Sume submit has returned a job id, a failed poll or result read should be retried against the same job, never answered by creating a new paid job. Retry 408, 425, 429, 500, 502, 503, 504 and 520 to 525, and treat a non-JSON body as a transport failure too.

This set comes from the repo's internal polling-resilience guidance rather than a public docs page. The public docs cover the error codes and the 409 behavior, and the sample keeps the two apart.

A classifier

It returns what to do next. It was run with four sample responses and printed retry, retry, check_status, ok.

The classifier is deliberately small. It does not parse the body for error codes, because the public error envelope is {error: {code, message, request_id, details}} and the right action depends on the code: for example 429 rate_limited and 429 queue_full both mean slow down, while 402 insufficient_credits and 400 invalid_request will not be fixed by retrying at all, and are not in the retry set. Put the finer code-based logic after this first split between transport trouble and real answers from the API, so each layer stays easy to test.

RETRYABLE = {408, 425, 429, 500, 502, 503, 504, 520, 522, 523, 524, 525}

def classify_readback(status: int, content_type: str, body: str) -> str:
    """Decide what to do with a poll or result read after a successful submit."""
    if status in RETRYABLE:
        return "retry"                      # same job id, never a new paid create
    if "json" not in content_type.lower():
        return "retry"                      # HTML error page from an edge, not Sume
    if status == 409:
        return "check_status"               # job_not_completed: read GET /v1/jobs/{id}
    if status == 404:
        return "stop"                       # unknown or foreign job id
    return "ok" if status < 300 else "stop"

print(classify_readback(524, "text/html", "<html>"))        # retry
print(classify_readback(200, "text/html", "<html>oops"))    # retry
print(classify_readback(409, "application/json", "{}"))     # wait
print(classify_readback(200, "application/json", "{}"))     # ok

The 409 trap

GET /v1/jobs/{id}/result answers 409 job_not_completed for any job that is not completed. That includes failed and canceled jobs, not only running ones. So a 409 does not mean wait; it means read GET /v1/jobs/{id} and look at the state. That is why the sample returns check_status.

Readback responses and the action (read 2026-10-06, Sume docs and polling guidance)
ResponseAction
408, 425, 429, 500, 502, 503, 504, 520 to 525Retry the same job id with backoff
200 with an HTML bodyRetry; an edge page, not Sume
409 job_not_completed on resultRead the job record, then decide
404Stop; the id is unknown to this key
2xx JSONUse it

Backoff and limits

Start at one to two seconds, double up to 30 seconds and add jitter. Stop only when the job is completed, failed or canceled, or when your own time budget is spent. A client timeout does not cancel the job, so running out of budget leaves the job running and billing unless you cancel it.

Tradeoffs

Retrying 429 politely means honoring the Retry-After or RateLimit headers when present. A wider retry set hides real outages for longer, so cap total time as well as attempts.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume