Webhook retry: fixed delay vs exponential backoff, with numbers
Fixed delay retries at even gaps; exponential backoff doubles them. Sume uses fixed for job webhooks and backoff for run webhooks: how long each keeps trying.

A fixed-delay retry waits the same time between attempts; exponential backoff doubles the wait each time, so it keeps trying for longer with the same attempt count. Sume uses both: job webhooks retry on a fixed delay of 30 seconds by default, and run webhooks use exponential backoff with jitter capped at one hour. Both stop after 10 attempts.
The difference matters most during an outage of your endpoint: the job schedule gives up in minutes, the run schedule in hours. The arithmetic below is derived from the documented formula and ignores jitter, Retry-After and the 10-second timeout each attempt can use.
What is the difference between fixed and exponential?
Fixed delay is simple and predictable, and it hits a struggling endpoint at a steady rate. Exponential backoff spaces attempts out so a down endpoint gets less traffic and more time to recover. Sume's Webhooks page says its job webhooks use "a fixed delay between attempts (30s by default), not exponential backoff".
How long does each Sume schedule keep trying?
Ten attempts leave nine gaps. For run webhooks the documented gap is min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h); counting attempts from 1, the gaps before the one-hour cap are 30, 60, 120, 240, 480, 960 and 1,920 seconds, and the last two are capped at 3,600.
| Job webhook | Run webhook | |
|---|---|---|
| Delay rule | Fixed, 30 s by default | 30 s doubling, jitter, 1 h cap |
| Gaps between 10 attempts | 9 x 30 s | 30, 60, 120, 240, 480, 960, 1,920, 3,600, 3,600 s |
| Total waiting | About 4.5 minutes | About 3 hours 3 minutes |
What happens when the window runs out?
Nothing happens to the work. Ten refused attempts leave a failed delivery and a job or run that still reached its real terminal state. Job docs say to keep status_url polling available for events that never arrive; for runs, fetch the receipt from result_url. Redeliver replays the real terminal event and does not consume one of the automatic 10.
Which schedule should my receiver plan for?
If your endpoint can be down for more than a few minutes, treat a job webhook as best effort and poll the job as your backup. For run webhooks, the longer window covers most short deploy gaps, and an endpoint that returns 429 or 503 with Retry-After can stretch a gap up to the one-hour cap. On both surfaces, return 2xx quickly after durably storing the event, and dedupe: job_id for jobs, request_id for runs.
Sources
Related posts
More in Developers
- Webhook unknown event type: return 204, not a 500
One Sume verifier covers run and job webhooks. Route on event and answer 204 for unknown types so a new event never becomes a 500 and a retry storm.
- Webhook signature mismatch: compare the secret fingerprint first
When a Sume webhook signature will not verify, compare x-sume-webhook-secret-fingerprint with the dashboard fingerprint. It is safe to paste into a ticket.
- Webhook secret rotation: no downtime with a 24-hour window
After you rotate a Sume webhook signing secret, deliveries carry two signatures for 24 hours. What the header looks like and how to redeploy safely.
- Webhook send test vs redeliver: which one replays a real job?
Send test posts a dummy webhook.test payload to a URL you type. Redeliver re-sends a real terminal event and does not use one of the automatic 10 attempts.
Written by Sume