Ten Sume job error categories and which ones are worth a retry

Sume failed jobs carry a category such as validation, quota or worker_timeout. Map all ten to retry, fix, top up or support in one small Python table.

5 min readSume
All posts

A failed Sume job carries public error metadata: a category, a stage, whether it is retryable, retry-after seconds, a public reason and a next action. The Sume docs list ten common categories. Four of them are safe to retry later with the same idempotency key (queue, generation_unavailable, runtime_unavailable, worker_timeout), two need you to poll first, and the rest need a change from you before any retry.

The ten categories

The table keeps the docs' guidance and adds one column for what a script should do. Prefer the job's own retryable and next_action fields over this table when they are present.

Job error categories and the action to take, as of 2026-10-08
CategoryDocs next actionScript action
validationCorrect the inputStop; fix the request
authExamine the API key and workspace accessStop; alert
quotaAdd funds, or decrease the request costStop; top up or shrink
queueRetry later with the same idempotency keyRetry later
generation_unavailableRetry laterRetry later
generation_rejectedExamine events; correct unsupported inputStop; read events
generation_timeoutPoll the status, or retry laterPoll first
runtime_unavailableRetry later; not aggressivelyRetry with a long delay
worker_timeoutPoll the status, or retry laterPoll first
internalExamine events; contact support with the request or job idStop; open a ticket

Why poll before you retry

A timeout category does not mean the work was lost. The job record may already have moved on, and a resubmit would create a second paid job. Read GET /v1/jobs/{id} first, then decide. When you do retry, send the original Idempotency-Key. The docs say not to retry unsafe submit requests without one.

The mapping in code

The function returns a short action word. Unknown categories fall through to support, so a new category never retries blindly.

ACTIONS = {
    "validation": "fix_input",
    "auth": "fix_auth",
    "quota": "top_up",
    "queue": "retry_later",
    "generation_unavailable": "retry_later",
    "generation_rejected": "read_events",
    "generation_timeout": "poll_first",
    "runtime_unavailable": "retry_slowly",
    "worker_timeout": "poll_first",
    "internal": "support",
}

def next_step(job_error: dict) -> str:
    if job_error.get("retryable") is False:
        return "stop"
    return ACTIONS.get(job_error.get("category", ""), "support")

print(next_step({"category": "worker_timeout"}))   # poll_first
print(next_step({"category": "quota"}))            # top_up
print(next_step({"category": "brand_new"}))        # support

Log the request id

Every error body has a request id that is safe to share with support. Do not include API keys, signed URLs, raw media URLs or private ids in a ticket.

Queue and capacity errors are not job errors

Do not mix the job categories with the submit errors. 402 insufficient_credits, 429 queue_full, 429 rate_limited and 503 provider_capacity_exceeded come back on the submit call, before a job exists. They follow their own rules: do not retry a 402 in a loop, wait for capacity on queue_full, honor retry-after on rate_limited, and retry later with the same key on provider_capacity_exceeded.

A job category tells you what happened to an accepted job. The two lists meet in one place, the idempotency key: when the guidance says retry later, retry with the same key, so that you cannot create a second paid job.

GET /v1/jobs/{id}/events returns a public timeline: job.created, job.queued, job.started, generation.submitted, and the terminal events. For generation_rejected and internal, read the events first. They show whether the job started and where it stopped, which is the information a support request needs along with the request id.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume