fal webhook retries: 31 attempts, 15 s timeout, vs Sume's 10

fal retries a failed webhook with backoff up to 31 times while the stored result lasts; Sume makes up to 10 attempts with a 10 second timeout. What to build.

4 min readSume
All posts

fal retries a failed webhook delivery with increasing backoff, up to 31 retries, until its stored result expires; the first attempt has a 15-second timeout and retries have 120 seconds. Sume makes up to 10 attempts in total with a 10-second timeout per attempt, then marks the delivery exhausted and leaves the job or run in its real state. Both want a fast 2xx, and both want your handler to tolerate repeats.

fal's numbers are from the Retry policy section of its Webhooks page, read on 2026-10-02. Sume's are from Webhooks and Run webhooks.

What is fal's retry policy?

Your endpoint must answer with a 2xx status. If a delivery times out, hits a network error, or gets a 4xx or 5xx, fal retries with increasing backoff until the stored result expires, which it puts at about 1 hour after the request completes, or about 6 minutes for results of 10 KB or more, up to a maximum of 31 retries.

fal tells you to design the handler to be idempotent and to expect repeat deliveries for the same request_id. If every attempt fails, you can usually still fetch the result from the queue while it is retained; results larger than 1 MB, and results for requests with payload storage disabled, are available only until the stored result expires.

What is Sume's retry policy?

Job webhooks (job.completed, job.failed, job.canceled) make up to 10 attempts total, with a fixed delay between attempts (30 seconds by default) rather than exponential backoff, and a 10-second timeout per attempt. Run webhooks (action.run.terminal, format.run.terminal, agent.run.terminal) also make up to 10 attempts, with a backoff of min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h), honouring Retry-After on 429 and 503.

Ten refused attempts leave a failed delivery, status exhausted, and a job that still reached its real terminal state. Sume says delivery is an optimization, never the only recovery path, and to keep polling status_url. Dedupe on job_id for jobs and request_id for runs.

How do the two compare?

The numbers differ, and so does what ends the retrying.

fal's Retry policy section and Sume's Webhooks and Run webhooks pages, read 2026-10-02.
ItemfalSume
SuccessAny 2xxAny 2xx
First attempt timeout15 seconds (retries: 120 seconds)10 seconds per attempt
Attempt cap31 retries10 attempts in total
What ends retriesStored result expires (about 1 hour, about 6 minutes at 10 KB or more) or the capThe cap; delivery then shows exhausted
SpacingIncreasing backoffJob webhooks: fixed, 30 s by default. Run webhooks: exponential with jitter, capped at 1 hour
Manual replayNot stated on the pageRedeliver per job or format run; does not use up the 10
Dedupe keyrequest_idjob_id (jobs), request_id (runs)

What should my receiver do?

Acknowledge fast after durably storing the event, and do the work afterwards. A slow endpoint burns the attempt budget on either platform. On Sume, a 10-second timeout is the tighter limit, so write the row, return 200, and process from a queue.

If deliveries stop arriving, fall back to reading the result directly. For Sume, POST /v1/jobs/{job_id}/webhook/redeliver (scope jobs:write) re-sends the real terminal event with a fresh timestamp and signature, and it works after the automatic attempts are exhausted.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume