Sume webhook retries: 10 attempts, 30 s apart, a 4.5 minute window
Sume retries a job webhook up to 10 times, 30 s apart by default. That is 270 s between first and last attempt, 370 s at worst. What to do after.

Sume tries a job webhook at most 10 times. The default delay between attempts is a fixed 30 seconds, so nine gaps add up to 270 seconds, or 4.5 minutes, from the first attempt to the last. If your endpoint is so slow that every attempt uses its full 10-second timeout, the worst case is 10 x 10 + 9 x 30 = 370 seconds. After that the delivery is marked failed, but the job itself keeps its real terminal state and you can still read it.
The numbers, with the arithmetic
The Sume docs give three delivery limits for job webhooks. The table turns them into windows you can plan around.
| Item | Value | How it is derived |
|---|---|---|
| Attempts | 10 total | Documented limit |
| Delay between attempts | 30 s by default, fixed | Documented; not exponential backoff |
| Timeout per attempt | 10 s | Documented; a slow endpoint uses the budget and is retried |
| Best-case window | 270 s | 9 gaps x 30 s |
| Worst-case window | 370 s | 10 attempts x 10 s + 9 gaps x 30 s |
| Retried on | Network errors and non-2xx | Documented |
What a receiver should do inside the window
Return any 2xx only after you have stored the event durably. Do the slow work after the response, not before it. Use job_id as your idempotency key, because the same terminal event can arrive more than once when an earlier attempt timed out on your side after you had already stored it.
Signature checks matter here. The verifier docs recommend rejecting callbacks whose timestamp is outside your replay tolerance, with five minutes as a reasonable default. The Sume docs say that Redeliver sends a fresh timestamp and signature. They do not say whether an automatic retry is re-signed. So log the reason for every rejected delivery in your first week, and check whether any 370-second tail ever lands outside your tolerance before you decide to widen it.
After the window: reconcile by polling
A failed delivery is not a failed job. The docs say delivery is an optimization and never the only recovery path, so keep the status_url poll available. The loop below reads the status of jobs whose webhook never arrived. If a job is terminal, handle it exactly as you would handle the webhook, keyed on job_id.
You can also ask Sume to send the real terminal event again with POST /v1/jobs/{job_id}/webhook/redeliver, which needs jobs:write. It works after all automatic attempts are used and does not consume one of the 10.
import asyncio, os
import httpx
ATTEMPTS, GAP, TIMEOUT = 10, 30, 10
print("best", (ATTEMPTS - 1) * GAP) # 270
print("worst", ATTEMPTS * TIMEOUT + (ATTEMPTS - 1) * GAP) # 370
async def reconcile(job_ids):
key = os.environ.get("SUME_API_KEY", "")
if not key:
raise SystemExit("set SUME_API_KEY")
headers = {"Authorization": f"Bearer {key}"}
async with httpx.AsyncClient(headers=headers, timeout=20) as c:
for jid in job_ids:
r = await c.get(f"https://api.sume.com/v1/jobs/{jid}/status")
r.raise_for_status()
s = r.json()
print(jid, s.get("terminal"), s.get("sume_status"))
asyncio.run(reconcile(["job_123"]))Checklist
- Plan for a 370-second tail, not 30 seconds.
- Dedupe on
job_id; never run a paid follow-up step twice. - Schedule a status sweep for any job with no event after about seven minutes.
- Use Redeliver, not a new submit, when you only lost the notification.
Make the receiver cheap to retry
Because the same terminal event can arrive up to ten times, the receiver should do very little before it answers. Verify the signature, write one row keyed on job_id with the raw body, and return 204. Move downloads, transcodes and database fan-out to a worker that reads the row. A handler that downloads a finished video before it answers can easily pass the 10-second timeout, and each timeout is one of your ten attempts.
The ten attempts also make a case for a unique constraint. If your insert fails because the row exists, treat it as success and still answer 2xx. Returning an error for a duplicate only produces more retries.
Two signals are worth an alert. The first is a job that has been terminal for longer than the worst-case window of 370 seconds without a stored event, which means delivery probably failed and the status sweep must run. The second is a rising number of rejected signatures or old timestamps, which can point at a clock problem or a rotation you forgot to deploy. Webhook delivery status, with the attempt count, is shown on the job object and in job events when available, so you can read it instead of guessing.
Sources
Related posts
More in Developers
- Sume webhook retries: 10 attempts, 30 s apart, 10 s timeout each
The delivery schedule for Sume job webhooks: 10 attempts, fixed 30 s spacing, 10 s timeout, about 4.5 minutes of retries, then redeliver and the status poll.
- Clock changes vs clock drift: what breaks Sume webhook checks
Sume webhook timestamps are Unix seconds, so time zone changes cannot break the 5-minute replay check; a drifting server clock can. A check to tell them apart.
- Sume webhook status is OK or ERROR; job status is completed or failed
A Sume job webhook body says status OK or ERROR, while the job endpoints say completed, failed or canceled. Branch on the event name and map both vocabularies.
- Swap the Sume video model with an env var and a catalog check in Node
Read the model id from VIDEO_MODEL, confirm it appears in GET /v1/videos/models, and fall back to sume/auto when it does not. A Node 20 script of 21 lines.
Written by Sume