Sume webhook retries: 10 attempts, 270 seconds of gaps, then poll

How long does Sume keep retrying a job webhook? Ten attempts, a 30 second default gap and a 10 second timeout, worked out in seconds, plus a Python dedupe.

5 min readSume
All posts

Sume tries a job webhook up to 10 times, with a fixed gap that is 30 seconds by default and a 10 second timeout on each attempt. Nine gaps of 30 seconds is 270 seconds, or 4 minutes 30 seconds, between the first attempt and the last. If your endpoint is down for longer than that, the delivery is exhausted and you must poll or redeliver.

The numbers come from the Delivery behavior section of the Webhooks page, read 2026-10-09. The page says the spacing is fixed, not exponential, and that a slow endpoint uses its 10 second budget and is retried like any other failure.

The arithmetic

Ten attempts leave nine gaps. The table shows the best and worst case with the default gap. Treat the second row as an upper bound for planning, not a promise, since the docs state the spacing as a default.

Retry window with default values, as of 2026-10-09, computed from the Webhooks page.
CaseCalculationResult
Endpoint refuses instantly9 gaps x 30 s270 s (4 min 30 s)
Every attempt hits the 10 s timeout9 x 30 s + 10 x 10 s370 s (6 min 10 s)
Attempts allowedFixed10 in total
SpacingFixed delay, not exponential30 s by default

What this means for a deploy

A rolling deploy that takes your receiver down for a minute is fine: the retries land after it comes back. A database migration that blocks writes for ten minutes is not, since all ten attempts will have been spent. Plan for the 270 second figure when you decide how long a receiver may be unavailable.

When the window is missed, the job still reached its real terminal state. The docs call delivery an optimization, not the only recovery path, and tell you to keep status_url polling available. A sweeper that lists recent jobs and compares them with the events you stored closes the gap without relying on redelivery.

Dedupe on job_id

Sume retries network errors and any non-2xx response, so a receiver that stores the event and then crashes before replying may see the same event again. The docs say to use job_id as the idempotency key. Store the event durably first, then return a 2xx. The snippet records each job_id in SQLite and tells the caller whether it is new.

import sqlite3

db = sqlite3.connect("webhooks.db")
db.execute("create table if not exists seen (job_id text primary key, body text)")


def first_delivery(job_id: str, body: str) -> bool:
    cur = db.execute(
        "insert or ignore into seen (job_id, body) values (?, ?)", (job_id, body)
    )
    db.commit()
    return cur.rowcount == 1


if __name__ == "__main__":
    print(first_delivery("job_123", "{}"))  # True
    print(first_delivery("job_123", "{}"))  # False: a retry or a redeliver

Redeliver is not a retry

A manual redeliver sends a fresh timestamp and signature and does not use one of the ten automatic attempts. Your dedupe table will see the same job_id again, so decide whether a redeliver should overwrite the stored body or be ignored. For a terminal-only event, ignoring it is usually right.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume