Sume webhook retry window: 4.5 minutes for jobs, 3 hours for runs
Job webhooks retry 10 times at a fixed 30 s, about 4.5 minutes. Run webhooks back off for about 3 hours. How long a deploy can take your receiver down.

A job webhook gives your receiver roughly 4.5 minutes to come back, and a Format, Action or Agent run webhook gives it roughly 3 hours. Both are 10 attempts, but the spacing differs: job webhooks use a fixed delay of 30 seconds, while run webhooks back off from 30 seconds and cap each wait at one hour. The totals below are my arithmetic from those documented numbers, not a figure Sume publishes, and they ignore the 10-second per-attempt timeout and jitter.
If you plan a deploy freeze, a database migration or a cloud-region failover around the receiver, the surface you call decides how long you can be dark before a delivery is marked failed. A rolling restart that takes 90 seconds is invisible to both. A 20-minute outage is invisible to run webhooks and fatal to job webhooks.
The arithmetic
Ten attempts leave nine gaps. For jobs that is 9 x 30 s = 270 s. For runs the documented schedule is min(max(30 s x 2^(attempt-1) with jitter, Retry-After), 1 h), which gives gaps of 30, 60, 120, 240, 480, 960, 1,920, then 3,600 and 3,600 seconds once the cap bites: 11,010 seconds, about 3 hours 3 minutes, before jitter.
The script below prints both. Run it as-is; it needs only Python 3.
| Surface | Attempts | Spacing | Total retry window (arithmetic) |
|---|---|---|---|
| Job webhook (job.completed / failed / canceled) | 10 | Fixed 30 s | about 4.5 min |
| Run webhook (*.run.terminal) | 10 | 30 s doubling, 1 h cap, honours Retry-After | about 3.1 h before jitter |
def gaps(attempts=10, base=30, cap=3600):
return [min(base * 2 ** (n - 1), cap) for n in range(1, attempts)]
job = [30] * 9
run = gaps()
for name, g in (("job", job), ("run", run)):
print(name, g, sum(g), "seconds", round(sum(g) / 3600, 2), "hours")What it means for a deploy
Treat the job window as the budget for any change that can take the receiver offline. If a migration holds the endpoint at 503 for ten minutes, every job webhook that terminated in that period burns its attempts and ends as a failed delivery. The job itself is fine: the job still reached its real terminal state, and the docs say delivery is an optimization, not the only recovery path.
So keep two recovery paths. First, a poller or sweeper that reads status_url for jobs that have been non-terminal longer than you expect. Second, redeliver: POST /v1/jobs/{job_id}/webhook/redeliver with a jobs:write key re-sends the real terminal event with a fresh timestamp and signature, works after the automatic attempts are exhausted, and does not consume one of the automatic 10. Run webhooks have their own redeliver on the run webhooks page.
- Shorter than 4 minutes: do nothing special, but still dedupe on
job_id. - Between 4 minutes and 3 hours: job webhooks need a sweep or redeliver; run webhooks recover by themselves.
- Longer than 3 hours: both need redeliver or a read of the run or job.
A planning checklist
Send a 503 with a Retry-After header from the maintenance path of a run-webhook receiver: the schedule uses the larger of its own backoff and your header, still capped at one hour. Job webhooks use a fixed delay, so a header does not stretch their window.
Acknowledge only after durably storing the event, and use job_id (jobs) or request_id (runs, equal to the run id) as the dedupe key, since a late redeliver can arrive after you already handled the event by polling. Verify the signature on the raw body first; the same verifier covers both event families, as covered in the redeliver post.
Measuring your own gap
You do not have to trust the arithmetic. On a staging receiver, return 503 for ten minutes and watch the delivery rows: job webhooks stop after their tenth attempt, while the same outage on a run webhook leaves the delivery in retrying. The webhook_delivery block on a run receipt, and the delivery status on a job, show the attempt count and last status code. Record how long your slowest deploy actually takes and compare it to the short window, not the long one.
Sources
Related posts
More in Developers
- GET /v1/jobs thread_id filter: why a teammate's job still 404s
The thread_id filter on the Sume jobs list narrows results and never widens what an API key can read. Teammates' jobs stay 404, and only the creator can cancel.
- jobs_result batch: read a wave when some jobs are still running
Sume's batch jobs_result returns one entry per id in request order. Read ok per entry, treat job_not_completed as running, and re-read only failed_job_ids.
- A kill switch for paid Sume submits: stop new jobs, cancel queued
Add an off switch to code that spends on the Sume API: check a flag before each submit, then cancel queued jobs; a 409 job_generation_already_started will bill.
- kling-3 on Sume: generate_audio true or false, and what changes
kling-3 has an audio toggle; MiniMax H3 and Omni audio is always on. The catalog lists separate audio-on and audio-off list rates.
Written by Sume