Sume webhook retries: 10 attempts, 30 s apart, 10 s timeout each

The delivery schedule for Sume job webhooks: 10 attempts, fixed 30 s spacing, 10 s timeout, about 4.5 minutes of retries, then redeliver and the status poll.

3 min readSume
All posts

Sume tries a job webhook up to 10 times, with a fixed delay between attempts that is 30 seconds by default, and gives each attempt 10 seconds to answer. If your endpoint is down, the attempts span roughly 4.5 minutes of spacing, so a deploy that takes longer than that loses the automatic deliveries.

Losing a delivery does not lose the job. The job still reaches its real terminal state, and the webhook docs name two ways to catch up: redeliver the event, or poll the status URL.

The schedule

Spacing is fixed, not exponential. The offsets below assume each failed attempt is followed by the default 30-second delay; a slow endpoint that burns its full 10 seconds pushes each later attempt back accordingly.

Default 30 s spacing, 10 attempts
AttemptEarliest start after the first
10 s
230 s
360 s
490 s
5120 s
6150 s
7180 s
8210 s
9240 s
10270 s

What to do inside the window

  • Acknowledge within 10 seconds: verify, store the event durably, return a 2xx, and do the real work afterwards.
  • Treat job_id as the idempotency key; an event can arrive twice.
  • Return a 2xx for event types you do not act on. Any non-2xx is retried, so an error for an unknown type would cause needless retries.

After attempt ten

The delivery status becomes exhausted while the job itself is unaffected. POST /v1/jobs/{job_id}/webhook/redeliver (scope jobs:write) re-sends the real terminal event with a fresh timestamp and signature, and it does not consume one of the automatic ten. "Send test" is different: it posts a dummy webhook.test payload to a URL you type and never replays a real job.

Keep a reconciler anyway. A cron that lists your stored job ids with no event and reads their status URL closes every gap without relying on any delivery.

Designing the receiver around the schedule

The practical consequence of a fixed 30-second delay is that short outages heal themselves and long ones do not. A rolling deploy of a few seconds loses nothing, because the next attempt arrives half a minute later. A database migration that keeps the endpoint down for ten minutes outlasts all ten attempts.

  • Put the endpoint behind a queue-backed handler so a slow downstream never counts against the 10-second budget.
  • Return 2xx only after the event is stored, so a crash between ack and store cannot lose it.
  • Alert on deliveries in retrying or exhausted, not only on your own error logs.
  • Schedule a reconciler that reads the status URL of every job older than a few minutes with no stored terminal event.

The delivery status vocabulary is pending, delivering, delivered, retrying, failed and exhausted, shown on the job object where it is available. Use it to see which of your events never landed before you redeliver.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume