Retry or not: a decision table for failed Sume submit calls

Which Sume errors deserve a retry with the same Idempotency-Key, which mean poll the job, and which mean fix the request. One table and a 15-line classifier.

4 min readSume
All posts

Retry with the same Idempotency-Key on 429 rate_limited, 429 queue_full and 503 provider_capacity_exceeded; poll the job on the two 409 job errors; and fix the request on 400, 401, 402, 409 idempotency_conflict, 413 and 415. Never retry an unsafe submit without a key.

The table below merges the errors page and the Generation admission page (both read 2026-10-10) into one decision per code.

The table

The last column is the one to code against. It has four values: retry with the same key, poll, fix, or alert a person.

Sume error codes and the client action (docs read 2026-10-10)
StatusCodeAction
400invalid_request, bad_requestFix the request, then submit again.
401unauthorizedStop every worker; fix the key.
402insufficient_creditsAdd funds or lower the cost; do not loop.
409job_not_completedPoll status; fetch the result when result_ready.
409job_generation_already_startedStop trying to cancel; the job completes or fails.
409idempotency_conflictFix the key or the payload; reuse a key only for an exact retry.
413 / 415payload_too_large, unsupported_media_typeShrink the body; send application/json.
429rate_limitedWait for retry-after, retry with the same key.
429queue_fullWait for jobs to finish or cancel queued ones; retry with the same key.
503provider_capacity_exceededRetry later with the same key.
503provider_not_configuredDo not retry aggressively; alert.

The classifier

It returns an action and a wait, and it treats retry-after as authoritative when present. Unknown codes fall through to an alert on purpose: a new code is a reason to look, not to retry.

RETRY_SAME_KEY = {"rate_limited", "queue_full", "provider_capacity_exceeded"}
STOP = {"unauthorized", "insufficient_credits", "invalid_request", "bad_request",
        "idempotency_conflict", "payload_too_large", "unsupported_media_type"}

def decide(status: int, code: str, retry_after: int | None = None):
    """Return (action, wait_seconds). 'poll' means read the job, never resubmit."""
    if code in RETRY_SAME_KEY:
        return ("retry_same_key", retry_after or 30)
    if code in ("job_not_completed", "job_generation_already_started"):
        return ("poll", retry_after or 5)
    if code == "provider_not_configured":
        return ("alert_do_not_loop", None)
    if code in STOP or status in (400, 401, 402, 409, 413, 415):
        return ("fix_request", None)
    return ("alert", None)

print(decide(429, "queue_full", 12), decide(402, "insufficient_credits"), decide(409, "job_not_completed"))

Two rules that sit above the table

First, a timeout is not an error code. If your own client gives up, you do not know whether the job exists. Resend with the same key and Sume returns the original job instead of billing a second one.

Second, queue_full and rate_limited are not the same event: one is accepted generation capacity, the other is request volume. The docs say full concurrency alone is not an error, because valid jobs are accepted as queued while queue capacity remains.

  • Cap retries and add jitter; the docs ask for backoff, not a number.
  • Log the request_id from every error body.
  • Page a human on 401 and 503 provider_not_configured.

Backoff shape

When the answer is yes, retry with exponential delay and random jitter, a hard attempt limit, and the retry-after header if it is present. The errors page says to use that header on a 429, and not to retry unsafe submit requests without an Idempotency-Key.

Use the same key on every retry of the same operation. A new key for each attempt turns a retry loop into several paid jobs. Cap the loop at a handful of tries and surface the failure after that, because an infinite loop on a 503 just keeps the pressure on the service that asked you to slow down.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume