Sume error retry matrix: retry with the same key, or fix the request

Which Sume API errors to retry with the same Idempotency-Key, which to fix first, and which to leave alone. A table and a small Python function that decides.

5 min readSume
All posts

Retry only when the error says the same request can succeed later, and then retry with the same Idempotency-Key. A 4xx at create means nothing ran and nothing was charged, so correct the call. The Formats error page also gives machine fields to use: retryable, retry_after_seconds and next_action, with values such as fix_input, add_funds, retry_later and poll_status.

The most frequent and costly mistake the docs call out is retrying a 403 insufficient_scope in a loop.

Retry decisions, as of 2026-10-09 (docs.sume.com/workflows/errors-and-credits)
StatusCodeAction
400invalid_requestFix the request; do not retry
401unauthorizedFix the key; check for two credentials
402insufficient_creditsAdd funds, or submit something cheaper
404not_found / model_not_foundCheck ids and the catalog
409job_not_completed / job_generation_already_startedPoll; cancel no longer possible
409idempotency_conflictDifferent payload on a used key: new key
429rate_limitedBack off; use retry-after
429queue_fullWait for capacity, then same key
503provider_capacity_exceededRetry later, same key
503provider_not_configuredDo not retry aggressively

A decision function

The function below encodes the table. It returns what to do, not whether to sleep, so the caller can decide how long to wait from retry-after.

RETRY_SAME_KEY = {
    ("429", "rate_limited"),
    ("429", "queue_full"),
    ("503", "provider_capacity_exceeded"),
}
POLL_INSTEAD = {("409", "job_not_completed"), ("409", "job_generation_already_started")}

def next_step(status: int, code: str) -> str:
    key = (str(status), code)
    if key in RETRY_SAME_KEY:
        return "retry_same_key"
    if key in POLL_INSTEAD:
        return "poll_status"
    if status == 402:
        return "add_funds"
    if status == 503:
        return "retry_later_slowly"
    return "fix_request"

Rules that go with it

Prefer error.retryable and next_action when the response has them, and fall back to this table when it does not. Cap your own retries and keep the delays growing. For rate limits, the response can include retry-after; use it when present.

  • Unsafe submits without an Idempotency-Key should not be retried at all.
  • queue_full is not a rate limit: wait for a terminal job or cancel queued ones.
  • A failed admission releases or refunds its reservation.
  • Log request_id on every non-2xx, so a support report is one search.

Backoff and limits

Pair next_step with a bounded sleep: honor retry-after when present, otherwise double the delay from one second up to a ceiling such as 60 seconds, and stop after a fixed number of tries. Never retry a paid submit that has no Idempotency-Key, because a lost response may hide a created job.

Alert when a code you treated as final appears often. A rise in 402 insufficient_credits is a balance problem; a rise in 429 queue_full means your waves are too large for the plan; a rise in 503 provider_capacity_exceeded is a reason to slow down and retry later.

Treat this as a habit, not a one-time fix. Write the rule down next to the code that calls the API, add a test that exercises it, and review it whenever the docs change. Check the linked documentation pages in the sources list for the current wording before you rely on any number here, because limits and field names can be revised, and a short test run costs far less than debugging a production incident.

When something does not match what you read here, capture the x-sume-request-id response header and the job or run id, and send those to support. Do not paste API keys, signing secrets or full request bodies into a ticket or a chat; the ids are enough for the team to find the request.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume