How long can your webhook be down? Sume job vs run retry windows
Sume job webhooks retry 10 times, 30 s apart: about 5 minutes. Run webhooks back off to roughly 3 hours. The arithmetic, and what Redeliver covers.

Two schedules, one signature
Sume has two webhook surfaces with the same signature scheme and different retry timing. Generation jobs (job.completed, job.failed, job.canceled) retry up to 10 attempts total with a fixed 30-second default gap. Runs of Actions, Formats and Agent Completions retry up to 10 attempts too, but on an exponential schedule.
That means your deploy window can be safe for one surface and fatal for the other. Work out the numbers once.
The arithmetic
Job webhooks: 10 attempts leave 9 gaps of 30 seconds, so 270 seconds between the first and last try. Each attempt may spend up to its 10-second timeout, which adds up to 100 seconds if your endpoint hangs, so plan for roughly 4.5 to 6 minutes.
Run webhooks: the gap is min(max(30 s times 2 to the power of attempt minus 1, with jitter, Retry-After), 1 hour). Without jitter and Retry-After the nine gaps are 30, 60, 120, 240, 480, 960, 1,920, 3,600 and 3,600 seconds, which is 11,010 seconds, about 3 hours. That is my sum from the documented formula, not a figure the docs state.
| Surface | Attempts | Gap rule | Window if every attempt fails |
|---|---|---|---|
| Job webhook | 10 | Fixed 30 s default | About 270 s plus attempt timeouts |
| Run webhook | 10 | 30 s doubling, cap 1 h | About 11,010 s before jitter |
After the window closes
When attempts run out you have a failed delivery and a job or run that still reached its real terminal state. Nothing is lost. Two tools recover it:
- Poll status_url or result_url; polling is the always-available path.
- Redeliver: POST /v1/jobs/{job_id}/webhook/redeliver (jobs:write) or POST /v1/format-runs/{run_id}/webhook/redeliver (formats:write). It sends the real terminal event with a fresh timestamp and signature, and it does not use one of the automatic 10.
- Dedupe on job_id for jobs and on request_id (equals run_id) for runs, because redelivery sends the same identity again.
Practical rule
Keep a deploy of your receiver under about four minutes and you survive job webhooks even in the worst case. Return 503 with a Retry-After during planned drains: run webhooks honor it. Either way, run a nightly reconciliation that polls anything still non-terminal after your longest expected render.
Sources
Related posts
More in Developers
- How long does AI lip sync take? fal says about a minute at 1080p
fal says a 1080p H3 Max lip-sync clip takes about a minute. On Sume the call is an async job, so poll with backoff. This Python example submits and polls.
- How long to wait for a Sume job webhook before you start polling
Job webhooks retry 10 times, 30 seconds apart, with a 10-second timeout each. Start your poll fallback at about 6 minutes, and here is the arithmetic.
- How many jobs_wait calls does a long video job need? 50 s slices
At the default 50 s slice a 10-minute render needs up to 12 jobs_wait calls, at the 55 s cap up to 11. Write the call budget into the agent instruction.
- How to cap a Sume Format run's spend: $500 ceiling, $400 default
Send generation_spend_cap_usd on each Format run. A value above 500 returns 400, a Format with no cap defaults to 400, and a run past its cap fails.
Written by Sume