Sume webhook retry window: 4.5 minutes for jobs, 3 hours for runs

Job webhooks retry 10 times at a fixed 30 s, about 4.5 minutes. Run webhooks back off for about 3 hours. How long a deploy can take your receiver down.

5 min readSume
All posts

A job webhook gives your receiver roughly 4.5 minutes to come back, and a Format, Action or Agent run webhook gives it roughly 3 hours. Both are 10 attempts, but the spacing differs: job webhooks use a fixed delay of 30 seconds, while run webhooks back off from 30 seconds and cap each wait at one hour. The totals below are my arithmetic from those documented numbers, not a figure Sume publishes, and they ignore the 10-second per-attempt timeout and jitter.

If you plan a deploy freeze, a database migration or a cloud-region failover around the receiver, the surface you call decides how long you can be dark before a delivery is marked failed. A rolling restart that takes 90 seconds is invisible to both. A 20-minute outage is invisible to run webhooks and fatal to job webhooks.

The arithmetic

Ten attempts leave nine gaps. For jobs that is 9 x 30 s = 270 s. For runs the documented schedule is min(max(30 s x 2^(attempt-1) with jitter, Retry-After), 1 h), which gives gaps of 30, 60, 120, 240, 480, 960, 1,920, then 3,600 and 3,600 seconds once the cap bites: 11,010 seconds, about 3 hours 3 minutes, before jitter.

The script below prints both. Run it as-is; it needs only Python 3.

read 2026-10-03
SurfaceAttemptsSpacingTotal retry window (arithmetic)
Job webhook (job.completed / failed / canceled)10Fixed 30 sabout 4.5 min
Run webhook (*.run.terminal)1030 s doubling, 1 h cap, honours Retry-Afterabout 3.1 h before jitter
def gaps(attempts=10, base=30, cap=3600):
    return [min(base * 2 ** (n - 1), cap) for n in range(1, attempts)]

job = [30] * 9
run = gaps()
for name, g in (("job", job), ("run", run)):
    print(name, g, sum(g), "seconds", round(sum(g) / 3600, 2), "hours")

What it means for a deploy

Treat the job window as the budget for any change that can take the receiver offline. If a migration holds the endpoint at 503 for ten minutes, every job webhook that terminated in that period burns its attempts and ends as a failed delivery. The job itself is fine: the job still reached its real terminal state, and the docs say delivery is an optimization, not the only recovery path.

So keep two recovery paths. First, a poller or sweeper that reads status_url for jobs that have been non-terminal longer than you expect. Second, redeliver: POST /v1/jobs/{job_id}/webhook/redeliver with a jobs:write key re-sends the real terminal event with a fresh timestamp and signature, works after the automatic attempts are exhausted, and does not consume one of the automatic 10. Run webhooks have their own redeliver on the run webhooks page.

  • Shorter than 4 minutes: do nothing special, but still dedupe on job_id.
  • Between 4 minutes and 3 hours: job webhooks need a sweep or redeliver; run webhooks recover by themselves.
  • Longer than 3 hours: both need redeliver or a read of the run or job.

A planning checklist

Send a 503 with a Retry-After header from the maintenance path of a run-webhook receiver: the schedule uses the larger of its own backoff and your header, still capped at one hour. Job webhooks use a fixed delay, so a header does not stretch their window.

Acknowledge only after durably storing the event, and use job_id (jobs) or request_id (runs, equal to the run id) as the dedupe key, since a late redeliver can arrive after you already handled the event by polling. Verify the signature on the raw body first; the same verifier covers both event families, as covered in the redeliver post.

Measuring your own gap

You do not have to trust the arithmetic. On a staging receiver, return 503 for ten minutes and watch the delivery rows: job webhooks stop after their tenth attempt, while the same outage on a run webhook leaves the delivery in retrying. The webhook_delivery block on a run receipt, and the delivery status on a job, show the attempt count and last status code. Record how long your slowest deploy actually takes and compare it to the short window, not the long one.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume