Sume webhook retry schedule: jobs every 30 s, runs exponential
Sume job webhooks retry up to 10 times about 30 seconds apart; run webhooks back off exponentially to one hour. What each gives your endpoint to recover.

Both Sume webhook surfaces make up to ten attempts with a 10-second timeout each, but they space attempts differently. Generation-job webhooks (job.completed, job.failed, job.canceled) wait a fixed delay between attempts, 30 seconds by default. Run webhooks (*.run.terminal) back off exponentially from 30 seconds and cap each wait at one hour. So an endpoint that is down for ten minutes can still receive a run webhook, but a job webhook will have given up.
The numbers come from Webhooks and Run webhooks, read 2026-10-02.
What are the two schedules?
Job webhooks: up to 10 attempts total, a fixed delay between attempts (30s by default), not exponential backoff, and a 10s timeout per attempt. Run webhooks: up to 10 attempts, with the wait before the next attempt equal to min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h).
How long does each schedule keep trying?
Ignoring jitter and the 10-second attempt time, nine waits separate ten attempts. The totals below are my arithmetic from the documented formulas, not a number Sume publishes.
| Surface | Wait after attempt 1 to 9 | Total waiting |
|---|---|---|
| Job webhook (default delay) | 30 s each, nine times | about 4.5 minutes |
| Run webhook | 30, 60, 120, 240, 480, 960, 1920 s, then 3600 s twice (capped) | about 3 hours |
What should my receiver do with that difference?
For job webhooks, treat the webhook as a hint and keep polling. Docs say delivery is an optimization, never the only recovery path: keep status_url polling available for events that never arrive. A deploy that takes your endpoint down for five minutes is enough to exhaust all ten job attempts.
For run webhooks, the longer window tolerates a deploy or an outage, but a delivery that ends failed or exhausted still leaves the run completed. Fetch it from result_url, fix the endpoint, and call redeliver. Redeliver does not consume one of the automatic ten.
In both cases return a 2xx after durably storing the event and do the work afterwards. A slow endpoint burns the 10-second budget and gets retried, so your handler may see the same event twice. Dedupe on job_id for jobs and request_id for runs.
Where do I see which attempt I am on?
Run receipts carry a webhook_delivery block with status, attempts, max_attempts (10), next_attempt_at, last_status_code and last_error. Job webhook status, including the attempt count, is visible on the job object and in job events when available. The status values are pending, delivering, delivered, retrying, failed and exhausted.
One more difference: a redirect is not followed on run webhooks, and a 3xx counts as a failed attempt, so register the final URL.
- Jobs: assume about 4.5 minutes of retry, then reconcile by polling.
- Runs: assume about 3 hours, then read
result_urland redeliver. - Both: answer 2xx fast, dedupe, and never rely on a webhook alone.
Can I lengthen the job webhook window?
Not from the API. The docs describe the 30-second delay as a default and say nothing about a per-request setting, so I would not build on changing it. The practical lever is on your side: keep the receiver trivial, queue the event, and run a periodic reconciliation that lists recent jobs and reads status_url for anything you have not seen a webhook for.
That reconciliation also covers the case where the delivery URL was valid at submit time and later stopped being reachable, since the URL is checked again at delivery.
Sources
Related posts
More in Developers
- jobs_result with job_ids: read ok per entry, re-read failed_job_ids
A Sume MCP jobs_result batch returns ok plus value or error per id. One job_not_completed does not fail the rest; re-read partial_failure.failed_job_ids.
- jobs_wait include_results: skip the second read, handle omitted ids
Set include_results true on Sume MCP jobs_wait: completed ids return jobs_result in results[]; ids that do not fit are named in results_omitted.job_ids.
- jobs_wait returned operator_stopped: what it means and what to do
operator_stopped on Sume jobs_wait: operations stopped the job. It is terminal, has no output, and its hold was refunded. Wait only on the pending ids.
- jobs_wait timeout_seconds 600 returns at 55 seconds, clamped
Sume MCP jobs_wait accepts timeout_seconds up to 600 but clamps it to 55 and says so in wait_slice_clamped. Repeat the wait instead.
Written by Sume