Sume job webhook retries: 370 s worst case, then poll or redeliver

Sume retries a job webhook 10 times, 30 s apart, with a 10 s timeout each: 270 s of gaps and up to 370 s in total. What to do when the window closes.

4 min readSume
All posts

A Sume job webhook makes at most 10 attempts, with a fixed delay of 30 seconds by default and a 10-second timeout on each attempt. That is 270 seconds of gaps when the endpoint refuses fast, and at most 370 seconds when every attempt runs to its timeout. After that the delivery is failed, but the job is not: it already reached its real terminal state.

The arithmetic

The Webhooks page lists the delivery behavior. Ten attempts means nine gaps between them. The spacing is fixed, not exponential.

Job webhook retry window (Sume docs, read 2026-10-09)
ScenarioAttempt timeGap timeTotal
Endpoint refuses at once (connection error or instant 5xx)about 0 s each9 x 30 s = 270 sabout 270 s
Endpoint hangs on every attempt10 x 10 s = 100 s9 x 30 s = 270 s370 s (6 min 10 s)

Slow handlers lose the race

A handler that takes 12 seconds to answer uses the whole 10-second budget, so Sume counts the attempt as failed and tries again, even if your work finished a moment later. The fix is the usual one: verify the signature, store the event durably, return a 2xx, and do the slow work afterward. The docs recommend job_id as your idempotency key. The durable-store post has a runnable version.

After the window closes

Three things remain available. First, keep polling status_url. The docs say that delivery is an optimization and never your only recovery path. Second, read the job object or its events, where the webhook delivery status and attempt count show when available. The delivery status values are pending, delivering, delivered, retrying, failed, and exhausted. Third, ask Sume to send the real terminal event again.

curl -X POST https://api.sume.com/v1/jobs/job_123/webhook/redeliver \
  -H "Authorization: Bearer $SUME_API_KEY"

Redeliver is not one of the ten

POST /v1/jobs/{job_id}/webhook/redeliver needs the jobs:write scope. It re-sends the terminal event of that job with a fresh timestamp and signature. It works after the automatic attempts are used up and does not count against the 10. It does not change the destination URL; a new URL is a new job. Send test, the other dashboard control, posts a dummy webhook.test payload and never replays a real job.

Choosing a recovery path

Redeliver is the fastest repair when your endpoint was down. It is a write request and needs the jobs:write scope, and it does not count against the 10 attempts. Send test posts a dummy webhook.test event so you can check your endpoint and signature code without a real job.

Polling is the fallback that needs no public endpoint. After the worst-case window, read GET /v1/jobs/{id}/status, stop on terminal: true, and fetch the result once result_ready is true. Do both: a webhook for speed, and a poll after the window for certainty.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume