Retry policy by Sume job error category

Sume failed jobs carry a category and retryability. Retry queue and capacity errors with a cap, stop on validation and quota, and skip blanket retries.

5 min readSume
All posts

A blanket "retry three times" wrapper wastes money on jobs that cannot succeed. Sume failed jobs expose public error metadata, including category, stage, retryability and a next action, per the errors guide. Use the category to choose: fix validation input, check auth, stop on quota, and retry queue and generation_unavailable later with the same idempotency key. Cap attempts, and log the category beside the job id.

Category to policy

The guide lists these common job error categories with a typical next action.

Job error categories and a retry policy (read 2026-10-04)
CategoryTypical next actionRetry?
validationFix inputNo, until the input changes
authCheck API key and workspace accessNo
quotaAdd funds or lower request costNo, until balance changes
queueRetry later with the same idempotency keyYes, with backoff
generation_unavailableRetry laterYes, with a cap

Retry the same job, or submit a new one

The docs tell you to reuse the same Idempotency-Key when a submit is retried after a client-side failure or a transient capacity error, so the retry cannot create a second job. That is different from re-running a job that already reached failed: replaying the key is meant to return the original job, so a deliberate new attempt on a failed job should use a new key, for example the old key plus an attempt number, and should be treated as a new paid submit. Keep the attempt count in your own record so each try is visible in your logs and ledger. For a transient capacity error, wait first.

Fall back to another model, carefully

If a pinned model keeps failing with generation_unavailable, a fallback to a second catalog id is reasonable for non-branded work. Limits differ per model, so check supported_durations and resolutions before you swap. Do not fall back on validation failures: the same bad input will fail again on the next model.

Do not retry what you cannot see

A client timeout does not cancel a job. The jobs guide says to poll with exponential backoff until completed, failed or canceled, and not to resubmit the original paid request just because a local process timed out. Poll the stored job id first; retry only when the job is terminal and failed, and cancel queued jobs you no longer want.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume