Sume webhook retry schedule: jobs every 30 s, runs exponential

Sume job webhooks retry up to 10 times about 30 seconds apart; run webhooks back off exponentially to one hour. What each gives your endpoint to recover.

5 min readSume
All posts

Both Sume webhook surfaces make up to ten attempts with a 10-second timeout each, but they space attempts differently. Generation-job webhooks (job.completed, job.failed, job.canceled) wait a fixed delay between attempts, 30 seconds by default. Run webhooks (*.run.terminal) back off exponentially from 30 seconds and cap each wait at one hour. So an endpoint that is down for ten minutes can still receive a run webhook, but a job webhook will have given up.

The numbers come from Webhooks and Run webhooks, read 2026-10-02.

What are the two schedules?

Job webhooks: up to 10 attempts total, a fixed delay between attempts (30s by default), not exponential backoff, and a 10s timeout per attempt. Run webhooks: up to 10 attempts, with the wait before the next attempt equal to min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h).

How long does each schedule keep trying?

Ignoring jitter and the 10-second attempt time, nine waits separate ten attempts. The totals below are my arithmetic from the documented formulas, not a number Sume publishes.

Approximate retry window, computed from the documented rules, read 2026-10-02
SurfaceWait after attempt 1 to 9Total waiting
Job webhook (default delay)30 s each, nine timesabout 4.5 minutes
Run webhook30, 60, 120, 240, 480, 960, 1920 s, then 3600 s twice (capped)about 3 hours

What should my receiver do with that difference?

For job webhooks, treat the webhook as a hint and keep polling. Docs say delivery is an optimization, never the only recovery path: keep status_url polling available for events that never arrive. A deploy that takes your endpoint down for five minutes is enough to exhaust all ten job attempts.

For run webhooks, the longer window tolerates a deploy or an outage, but a delivery that ends failed or exhausted still leaves the run completed. Fetch it from result_url, fix the endpoint, and call redeliver. Redeliver does not consume one of the automatic ten.

In both cases return a 2xx after durably storing the event and do the work afterwards. A slow endpoint burns the 10-second budget and gets retried, so your handler may see the same event twice. Dedupe on job_id for jobs and request_id for runs.

Where do I see which attempt I am on?

Run receipts carry a webhook_delivery block with status, attempts, max_attempts (10), next_attempt_at, last_status_code and last_error. Job webhook status, including the attempt count, is visible on the job object and in job events when available. The status values are pending, delivering, delivered, retrying, failed and exhausted.

One more difference: a redirect is not followed on run webhooks, and a 3xx counts as a failed attempt, so register the final URL.

  • Jobs: assume about 4.5 minutes of retry, then reconcile by polling.
  • Runs: assume about 3 hours, then read result_url and redeliver.
  • Both: answer 2xx fast, dedupe, and never rely on a webhook alone.

Can I lengthen the job webhook window?

Not from the API. The docs describe the 30-second delay as a default and say nothing about a per-request setting, so I would not build on changing it. The practical lever is on your side: keep the receiver trivial, queue the event, and run a periodic reconciliation that lists recent jobs and reads status_url for anything you have not seen a webhook for.

That reconciliation also covers the case where the delivery URL was valid at submit time and later stopped being reachable, since the URL is checked again at delivery.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume