Retry policy by Sume job error category
Sume failed jobs carry a category and retryability. Retry queue and capacity errors with a cap, stop on validation and quota, and skip blanket retries.

A blanket "retry three times" wrapper wastes money on jobs that cannot succeed. Sume failed jobs expose public error metadata, including category, stage, retryability and a next action, per the errors guide. Use the category to choose: fix validation input, check auth, stop on quota, and retry queue and generation_unavailable later with the same idempotency key. Cap attempts, and log the category beside the job id.
Category to policy
The guide lists these common job error categories with a typical next action.
| Category | Typical next action | Retry? |
|---|---|---|
| validation | Fix input | No, until the input changes |
| auth | Check API key and workspace access | No |
| quota | Add funds or lower request cost | No, until balance changes |
| queue | Retry later with the same idempotency key | Yes, with backoff |
| generation_unavailable | Retry later | Yes, with a cap |
Retry the same job, or submit a new one
The docs tell you to reuse the same Idempotency-Key when a submit is retried after a client-side failure or a transient capacity error, so the retry cannot create a second job. That is different from re-running a job that already reached failed: replaying the key is meant to return the original job, so a deliberate new attempt on a failed job should use a new key, for example the old key plus an attempt number, and should be treated as a new paid submit. Keep the attempt count in your own record so each try is visible in your logs and ledger. For a transient capacity error, wait first.
Fall back to another model, carefully
If a pinned model keeps failing with generation_unavailable, a fallback to a second catalog id is reasonable for non-branded work. Limits differ per model, so check supported_durations and resolutions before you swap. Do not fall back on validation failures: the same bad input will fail again on the next model.
Do not retry what you cannot see
A client timeout does not cancel a job. The jobs guide says to poll with exponential backoff until completed, failed or canceled, and not to resubmit the original paid request just because a local process timed out. Poll the stored job id first; retry only when the job is terminal and failed, and cancel queued jobs you no longer want.
Sources
Related posts
More in Developers
- Expiring API keys: Sume key metadata and rotation habits
OpenAI added enforced key lifetimes in Sep 2026. Sume's docs list key id, name, prefix, scopes and last-used time, so rotate on a schedule you keep.
- Run an OpenRouter-style image script against Sume: four changes
Moving a script from OpenRouter's /api/v1/images to Sume's /v1/images: base64 becomes a hosted URL, 202 jobs appear, stream and seed return 400. Python check.
- Run your own 10-clip word error test on Sume STT for 10 cents
Microsoft ranks MAI-Transcribe-2-Streaming first on Artificial Analysis. To know your own audio, score 10 one-minute clips with a 15-line WER function.
- Same prompt, four Sume video models: a Python script that logs cost
Submit one prompt to wan-3.0, minimax-h3, minimax-h3-max and seedance-2.5 on /v1/videos, poll each job, and print usage.cost per clip. Runnable as written.
Written by Sume