Sume webhook retries: 10 attempts, 30 s apart, 10 s timeout each
The delivery schedule for Sume job webhooks: 10 attempts, fixed 30 s spacing, 10 s timeout, about 4.5 minutes of retries, then redeliver and the status poll.

Sume tries a job webhook up to 10 times, with a fixed delay between attempts that is 30 seconds by default, and gives each attempt 10 seconds to answer. If your endpoint is down, the attempts span roughly 4.5 minutes of spacing, so a deploy that takes longer than that loses the automatic deliveries.
Losing a delivery does not lose the job. The job still reaches its real terminal state, and the webhook docs name two ways to catch up: redeliver the event, or poll the status URL.
The schedule
Spacing is fixed, not exponential. The offsets below assume each failed attempt is followed by the default 30-second delay; a slow endpoint that burns its full 10 seconds pushes each later attempt back accordingly.
| Attempt | Earliest start after the first |
|---|---|
| 1 | 0 s |
| 2 | 30 s |
| 3 | 60 s |
| 4 | 90 s |
| 5 | 120 s |
| 6 | 150 s |
| 7 | 180 s |
| 8 | 210 s |
| 9 | 240 s |
| 10 | 270 s |
What to do inside the window
- Acknowledge within 10 seconds: verify, store the event durably, return a 2xx, and do the real work afterwards.
- Treat
job_idas the idempotency key; an event can arrive twice. - Return a 2xx for event types you do not act on. Any non-2xx is retried, so an error for an unknown type would cause needless retries.
After attempt ten
The delivery status becomes exhausted while the job itself is unaffected. POST /v1/jobs/{job_id}/webhook/redeliver (scope jobs:write) re-sends the real terminal event with a fresh timestamp and signature, and it does not consume one of the automatic ten. "Send test" is different: it posts a dummy webhook.test payload to a URL you type and never replays a real job.
Keep a reconciler anyway. A cron that lists your stored job ids with no event and reads their status URL closes every gap without relying on any delivery.
Designing the receiver around the schedule
The practical consequence of a fixed 30-second delay is that short outages heal themselves and long ones do not. A rolling deploy of a few seconds loses nothing, because the next attempt arrives half a minute later. A database migration that keeps the endpoint down for ten minutes outlasts all ten attempts.
- Put the endpoint behind a queue-backed handler so a slow downstream never counts against the 10-second budget.
- Return 2xx only after the event is stored, so a crash between ack and store cannot lose it.
- Alert on deliveries in
retryingorexhausted, not only on your own error logs. - Schedule a reconciler that reads the status URL of every job older than a few minutes with no stored terminal event.
The delivery status vocabulary is pending, delivering, delivered, retrying, failed and exhausted, shown on the job object where it is available. Use it to see which of your events never landed before you redeliver.
Sources
Related posts
More in Developers
- Swift: URLSession async/await for one 30-second Wan 3.0 job
A 28-line main.swift that submits wan-3.0 for 30 seconds, polls with Task.sleep and saves the MP4. Runs on macOS or Linux with swiftc.
- Switch video models by changing one string: what can still break
On Sume's /v1/videos you swap the model id and keep the body. Duration range, resolution and aspect ratio are the three fields that may need adjusting.
- sync, subscribe, async or webhook: which Sume mode for a video job
Sume's sync and subscribe modes wait at most 30 seconds, then return a job id. Use async or webhook for video; Python that survives a timed-out wait.
- rate_limited or queue_full? One Python submit handler for both 429s
Two different 429s need two different waits. A Python handler reads error.code, sleeps on retry-after for rate_limited, and waits for capacity on queue_full.
Written by Sume