fal retries failed requests up to 10 times; on Sume the retry is yours
fal's queue retries transient errors up to 10 times unless you send X-Fal-No-Retry. On Sume the retry is yours: reuse the Idempotency-Key, never resubmit blind.

fal's queue documentation says failed requests retry automatically up to 10 times across transient errors, and that you can turn this off with the X-Fal-No-Retry header. Sume's public API makes no such server-side promise for your submit: the retry is the caller's decision, and the safe way to make it is to repeat the same submit with the same Idempotency-Key. A retry with the same key returns the original job instead of billing a second one.
Where the two approaches differ
On a queue that retries for you, the question you ask is how to switch retries off for work that must not repeat. On Sume the question is how to retry exactly once per intent. The answer is a key per paid intent, such as one key per clip in a batch, and a rule that a timeout in your own process is never a reason to submit again. Poll the job you already have.
| Topic | fal queue | Sume |
|---|---|---|
| Automatic retry of failed requests | Up to 10 times on transient errors | Not promised for your submit |
| Opt-out | X-Fal-No-Retry header | Not applicable |
| Safe client retry | Not described on the page | Same Idempotency-Key returns the original job |
| Capacity error | Requests are never dropped, per the page | 429 queue_full or 503 provider_capacity_exceeded: retry later, same key |
The error table to follow
Sume's error docs give an action for each class. For 429 rate_limited, back off and obey retry-after. For queue_full, wait for a queued or processing job to finish or cancel one, then retry with the same key. For provider_capacity_exceeded, retry later with the same key. For provider_not_configured or runtime_unavailable, do not retry aggressively. A failed job also carries public error metadata (category, stage, retryability and retry-after seconds), so a retry loop can read the answer instead of guessing.
- Reuse the key only for an exact retry. A different payload under the same key is
409 idempotency_conflict. - Cap your own retries, for example three, and then surface the error.
- If a local timeout fired, call
GET /v1/jobs/{id}/status. Do not resubmit. - Do not retry unsafe submits without a key. The docs say so for
429too.
A migration note
Teams moving a queue worker from fal to Sume often carry over an assumption that the platform will retry. Add the retry loop on your side, put the idempotency key in the first line of the submit helper, and write one test that submits twice with the same key and asserts a single job id comes back. That one test catches most double-billing bugs before they ship.
Sources
Related posts
More in Developers
- fal never drops queued requests; Sume can answer 429 queue_full
fal's queue docs say queued requests are never dropped. Sume caps accepted jobs per plan and returns 429 queue_full when full. Plan for the gap.
- Fallback chain for a 30-second AI video: first catalog row that fits
Kling 4.0 is rolling out in stages. Pick the first Sume video id whose catalog row lists 30 seconds, and stop with an error if none does.
- Fan out one Sume clip to three platforms: one webhook, three task keys
Receive one job.completed webhook per Sume job, then enqueue a task per platform. Key each task by job id plus platform so a retry never double-posts.
- Sume video content 409: job_not_completed vs job_failed in Python
A 409 from /v1/videos/{id}/content means two things. job_not_completed is retryable, job_failed is not. Here is a Python handler that tells them apart.
Written by Sume