Sume webhook retry schedule: 30s doubling, 10 tries, 1h cap

Sume Format run webhooks retry up to 10 times with min(max(30s x 2^(attempt-1), Retry-After), 1h). The arithmetic of that schedule and what to do when it ends.

5 min readSume
All posts

Sume tries a Format run webhook up to 10 times. The documented backoff is min(max(30s × 2^(attempt−1) with jitter, Retry-After), 1h), each attempt has a 10-second timeout, and a redirect counts as a failed attempt. Ignoring jitter, the nine waits add up to about three hours, after which webhook_delivery.status becomes exhausted and the run is still completed.

The formula and limits are from Run webhooks; the status values are on Runs and results. The schedule below is my arithmetic from that formula, not a promise of exact times.

What is the retry schedule in numbers?

Taking the formula at face value, with no jitter and no Retry-After, and numbering each wait by the failed attempt it follows:

Computed backoff from the documented formula, jitter and attempt time ignored (read 2026-10-03)
After failed attemptWaitElapsed in waits
130 s30 s
260 s1 min 30 s
3120 s3 min 30 s
4240 s7 min 30 s
5480 s15 min 30 s
6960 s31 min 30 s
71920 s63 min 30 s
83600 s (capped)123 min 30 s
93600 s (capped)183 min 30 s

How does Retry-After change it?

The wait is the longer of the exponential figure and your Retry-After, capped at one hour. Sume honours it on 429 and 503 answers from your endpoint. A Retry-After: 5 can never shorten the schedule, because the exponential value wins when it is larger; a Retry-After: 1800 early on stretches that wait to 30 minutes.

The practical rule: if you are shedding load, answer 503 with a real Retry-After; do not answer 200 and drop the work, because then no retry comes.

What counts as a failed attempt?

Anything but a 2xx within 10 seconds. The docs name three cases: a timeout, a 3xx (redirects are not followed, so register the final URL), and a URL that fails re-validation at delivery time. The URL must be public HTTPS; localhost, private ranges, credentials in the URL and plain HTTP are 400 invalid_request at create.

last_status_code and last_error on the receipt's webhook_delivery block say which one happened. last_error is Sume's transport error, never your response body.

What do I do when the schedule runs out?

Nothing about the run changes. Read the receipt from result_url, fix the endpoint, then replay with POST /v1/format-runs/{run_id}/webhook/redeliver (needs formats:write, empty body). The docs say redelivery does not consume one of the automatic ten and still works after exhaustion.

Do both in production: take the webhook as the fast path and keep a poller on result_url as the backup for the day your endpoint is down.

How should the receiver be built around this?

Record the event durably, answer 2xx, then do the work. A slow endpoint burns the 10-second attempt budget and gets retried while it is still working, and every retry carries the same request_id, so dedupe on that.

Check the webhook_delivery block rather than assuming. It reports status, attempts, max_attempts, next_attempt_at, last_attempt_at, last_status_code and last_error, so you can see where in the schedule a run is.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume