Format webhook: 10 s timeout, 10 attempts, about 3 h 5 min of retries

A slow Format webhook gets 10 tries. Without jitter, the last starts 11,010 s after the first; with 10 s timeouts the span is about 3 h 5 min. Code included.

5 min readSume
All posts

If your endpoint takes longer than 10 seconds to answer, Sume counts the attempt as failed and tries again, up to 10 attempts in total. By the docs' formula, and ignoring jitter and any Retry-After header, the tenth attempt starts 11,010 seconds after the first, which is 3 hours and 3.5 minutes. If every attempt also runs the full 10-second timeout, the delivery window ends about 3 hours 5 minutes after the first attempt began.

The practical rule: write the event to durable storage, return a 2xx, and do the slow work afterwards.

The retry schedule

The docs give backoff as min(max(30 s x 2^(attempt-1) with jitter, Retry-After), 1 h). Taking the 30-second base without jitter, the wait before each retry and the running total are below.

Sume run webhook retry arithmetic from the docs formula, read 2026-10-09, jitter and Retry-After ignored
AttemptWait before it (s)Seconds since attempt 1
1none0
23030
36090
4120210
5240450
6480930
79601890
819203810
936007410
10360011010

What the numbers mean for your receiver

The same request_id arrives each time, so an at-least-once delivery is safe if your store ignores repeats.

  • Success is any 2xx. A 3xx counts as a failed attempt because Sume does not follow redirects.
  • The timeout is 10 seconds per attempt, so a handler that renders a thumbnail or calls another API before replying will time out and be retried.
  • After 10 attempts the receipt's webhook_delivery.status is exhausted; the run itself is unaffected and you can still poll result_url.
  • Dedupe on the envelope request_id, which equals the run id and is stable across retries.

A receiver that replies first

This stdlib Python server verifies the HMAC-SHA256 signature over the timestamp, a dot and the raw body, rejects timestamps more than 300 seconds old, and refuses to start without a secret. It answers 204 before any further work.

import hashlib, hmac, os, time
from http.server import BaseHTTPRequestHandler, HTTPServer

SECRET = os.environ.get("SUME_COM_WEBHOOK_SIGNING_SECRET", "")
if not SECRET:
    raise SystemExit("set SUME_COM_WEBHOOK_SIGNING_SECRET")

class Hook(BaseHTTPRequestHandler):
    def do_POST(self):
        raw = self.rfile.read(int(self.headers.get("content-length", 0)))
        ts = self.headers.get("x-sume-webhook-timestamp", "")
        sig = self.headers.get("x-sume-webhook-signature", "")
        mac = hmac.new(SECRET.encode(), ts.encode() + b"." + raw, hashlib.sha256)
        want = "sume-v1=" + mac.hexdigest()
        fresh = ts.isdigit() and abs(time.time() - int(ts)) <= 300
        ok = fresh and hmac.compare_digest(sig, want)
        self.send_response(204 if ok else 401)
        self.end_headers()
        if ok:
            # Save raw to a durable queue here; do slow work after replying.
            print("accepted", len(raw), "bytes")

HTTPServer(("", 8080), Hook).serve_forever()

Planning around the window

Treat 3 hours as the outer edge for an endpoint that is down, not a service level. If your receiver is down for a deploy of ten minutes, the fourth attempt, at 210 seconds, and the fifth, at 450 seconds, still land in the outage, but the sixth, at 930 seconds (15.5 minutes), should succeed. A longer outage uses more attempts, and one that outlasts 11,010 seconds loses the push entirely.

Keep a poll as a backstop. A run exposes status_url and result_url, and the docs say the webhook is the alternative to a poll loop, not a replacement for being able to poll. A sweep every hour that lists recent runs and fetches any you have not recorded closes the gap without a loop per run.

A handler that does its work before replying holds the connection for the whole job. If that work takes 12 seconds, every attempt times out at 10 seconds and Sume retries, even though your code succeeded each time. You then process the same event several times, which is why the request_id check matters. Write the raw body and headers to storage first, answer, and let a worker read the queue.

A run whose delivery was exhausted still has a result. Read it from result_url, or ask Sume to send the event again with POST /v1/format-runs/{run_id}/webhook/redeliver. Canceled and skipped runs never send a webhook, so do not count them as failed deliveries.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume