How long can your webhook be down? Sume job vs run retry windows

Sume job webhooks retry 10 times, 30 s apart: about 5 minutes. Run webhooks back off to roughly 3 hours. The arithmetic, and what Redeliver covers.

4 min readSume
All posts

Two schedules, one signature

Sume has two webhook surfaces with the same signature scheme and different retry timing. Generation jobs (job.completed, job.failed, job.canceled) retry up to 10 attempts total with a fixed 30-second default gap. Runs of Actions, Formats and Agent Completions retry up to 10 attempts too, but on an exponential schedule.

That means your deploy window can be safe for one surface and fatal for the other. Work out the numbers once.

The arithmetic

Job webhooks: 10 attempts leave 9 gaps of 30 seconds, so 270 seconds between the first and last try. Each attempt may spend up to its 10-second timeout, which adds up to 100 seconds if your endpoint hangs, so plan for roughly 4.5 to 6 minutes.

Run webhooks: the gap is min(max(30 s times 2 to the power of attempt minus 1, with jitter, Retry-After), 1 hour). Without jitter and Retry-After the nine gaps are 30, 60, 120, 240, 480, 960, 1,920, 3,600 and 3,600 seconds, which is 11,010 seconds, about 3 hours. That is my sum from the documented formula, not a figure the docs state.

Retry windows computed from the Sume webhook docs, read 2026-10-05
SurfaceAttemptsGap ruleWindow if every attempt fails
Job webhook10Fixed 30 s defaultAbout 270 s plus attempt timeouts
Run webhook1030 s doubling, cap 1 hAbout 11,010 s before jitter

After the window closes

When attempts run out you have a failed delivery and a job or run that still reached its real terminal state. Nothing is lost. Two tools recover it:

  • Poll status_url or result_url; polling is the always-available path.
  • Redeliver: POST /v1/jobs/{job_id}/webhook/redeliver (jobs:write) or POST /v1/format-runs/{run_id}/webhook/redeliver (formats:write). It sends the real terminal event with a fresh timestamp and signature, and it does not use one of the automatic 10.
  • Dedupe on job_id for jobs and on request_id (equals run_id) for runs, because redelivery sends the same identity again.

Practical rule

Keep a deploy of your receiver under about four minutes and you survive job webhooks even in the worst case. Return 503 with a Retry-After during planned drains: run webhooks honor it. Either way, run a nightly reconciliation that polls anything still non-terminal after your longest expected render.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume