Sume webhook retries: 10 attempts 30 seconds apart, size your downtime

Sume retries a failing webhook up to 10 attempts at a default 30-second spacing, about 4.5 minutes. Past that, redeliver or poll after a publisher outage.

5 min readSume
All posts

Sume retries a failed job webhook up to 10 attempts in total, with a fixed delay between attempts that is 30 seconds by default and a 10-second timeout for each attempt. At default spacing that is nine gaps of 30 seconds, so a publisher outage longer than about 4.5 minutes can lose the automatic deliveries; the jobs still finish.

After the attempts are used, recover with POST /v1/jobs/{job_id}/webhook/redeliver (jobs:write) or by polling status_url. Redeliver sends a fresh timestamp and signature and does not use one of the automatic 10.

What are the delivery rules?

From Sume's webhook docs.

Sume webhook delivery behavior (read 2026-10-05)
RuleValue
AttemptsUp to 10 in total
SpacingFixed, 30 seconds by default, not exponential backoff
Timeout10 seconds for each attempt
SuccessAny 2xx after you store the event durably
After exhaustionRedeliver per job, or poll status_url

How do I find the jobs I missed?

After an outage, list the job ids you submitted that have no stored terminal event, read each status, and redeliver or fetch the result. The sample finds the gaps.

def missing(submitted: list, received: set) -> list:
    return [job_id for job_id in submitted if job_id not in received]


if __name__ == "__main__":
    print(missing(["job_1", "job_2", "job_3"], {"job_1", "job_3"}))

How long can I be down?

Plan for the arithmetic: the 4.5 minutes is a calculation from the documented defaults, not a promise, so treat any deploy that takes the receiver down longer as needing a sweep. The post on a reconcile sweeper shows the loop.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume