Failed Sume job: read category, retryable and next_action first
Failed jobs expose public error metadata: category, stage, retryable, retry-after, reason and next action. Map each category to retry, fix or stop.

A failed Sume job carries public error metadata: category, stage, retryability, retry-after seconds, a public reason and a next action. Read those before deciding to retry. Validation and quota failures need a fix on your side, while queue, unavailable and timeout failures are candidates to retry later with the same idempotency key.
Where to read it
GET /v1/jobs/{id} returns the job record with its error. GET /v1/jobs/{id}/result is only for completed jobs and answers 409 job_not_completed for others, so do not read failures from it. On /v1/videos, the poll response's error is the same public remap. Provider payloads are never exposed as public fields.
Category to action
The errors guide lists the common categories and their typical next action.
Use the table as a default policy, then let retryable and retry-after override it.
| Category | Typical next action | Retry? |
|---|---|---|
| validation | Correct the input | No |
| auth | Check the API key and workspace access | No |
| quota | Add funds or reduce the cost | No |
| queue | Retry later, same idempotency key | Yes |
| generation_unavailable | Retry later | Yes |
| generation_rejected | Read events and fix the unsupported input | No |
| generation_timeout | Poll status or retry later | Yes |
| worker_timeout | Poll status or retry later | Yes |
| internal | Read events and contact support with the job id | No |
A small decision function
Treat the server's retryable flag as the first signal and fall back to the category.
const RETRY = new Set([
'queue', 'generation_unavailable', 'generation_timeout',
'worker_timeout', 'runtime_unavailable',
]);
export function shouldRetry(err) {
if (typeof err.retryable === 'boolean') return err.retryable;
return RETRY.has(err.category);
}Retrying without double billing
Resubmit with the same Idempotency-Key only when the body is identical. If you changed the prompt or inputs, that is a new intent and needs a new key. Failed and canceled image generations are not charged, and the job events (job.failed, webhook.delivery) give a public timeline when you need to debug.
Related posts
More in Developers
- fal retries failed requests up to 10 times; on Sume the retry is yours
fal's queue retries transient errors up to 10 times unless you send X-Fal-No-Retry. On Sume the retry is yours: reuse the Idempotency-Key, never resubmit blind.
- fal never drops queued requests; Sume can answer 429 queue_full
fal's queue docs say queued requests are never dropped. Sume caps accepted jobs per plan and returns 429 queue_full when full. Plan for the gap.
- Fallback chain for a 30-second AI video: first catalog row that fits
Kling 4.0 is rolling out in stages. Pick the first Sume video id whose catalog row lists 30 seconds, and stop with an error if none does.
- Fan out one Sume clip to three platforms: one webhook, three task keys
Receive one job.completed webhook per Sume job, then enqueue a task per platform. Key each task by job id plus platform so a retry never double-posts.
Written by Sume