Webhook returns 503 with Retry-After: how Sume spaces the retries

If your webhook endpoint answers 429 or 503 with Retry-After, Sume run webhooks wait at least that long, capped at one hour. Here is how to shed load safely.

5 min readSume
All posts

When a Sume run webhook reaches an endpoint that answers 429 or 503 with a Retry-After header, Sume waits at least that long before the next attempt, up to a cap of one hour. The documented wait is min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h), so a short Retry-After never makes Sume retry sooner than its own backoff.

That makes a temporary 503 a legitimate way to shed load during a spike, as long as you remember the budget is still ten attempts. Sources: Run webhooks and Runs and results, read 2026-10-02.

Does Retry-After apply to job webhooks too?

The docs I read describe Retry-After handling only for run webhooks. The generation-job page describes a fixed delay between attempts, 30 seconds by default, and does not mention Retry-After. If you need a pause on a job webhook, do not depend on the header; accept the event, store it and process later.

What counts as a failed attempt?

Anything that is not a timely 2xx uses up an attempt.

What your endpoint returns and what Sume does, run webhooks, read 2026-10-02
Your endpointResult
2xx within 10 secondsDelivered
429 or 503 with Retry-AfterFailed attempt; next wait is the longer of the backoff and Retry-After, capped at 1 hour
Slow response over 10 secondsTimed out; retried
3xx redirectNot followed; counts as a failed attempt
Other non-2xx or network errorFailed attempt; retried until ten are used

How do I shed load without dropping events?

The safest pattern is to verify, write the raw event to durable storage, and answer 204 immediately. Only answer 503 with Retry-After when you cannot even store it. Below is a minimal stdlib receiver that refuses to start without a secret, checks the signature and replay window, and sheds load when told to. During a secret rotation the signature header carries several comma-separated entries, so it accepts any sume-v1= match.

import hashlib, hmac, os, time
from http.server import BaseHTTPRequestHandler, HTTPServer

SECRET = os.environ.get("SUME_COM_WEBHOOK_SIGNING_SECRET", "")
if not SECRET:
    raise SystemExit("set SUME_COM_WEBHOOK_SIGNING_SECRET")
BUSY = os.environ.get("BUSY") == "1"

def verify(body, ts, header):
    if not ts.isdigit() or abs(time.time() - int(ts)) > 300:
        return False
    mac = hmac.new(SECRET.encode(), ts.encode() + b"." + body, hashlib.sha256)
    want = "sume-v1=" + mac.hexdigest()
    return any(hmac.compare_digest(e.strip(), want) for e in header.split(","))

class H(BaseHTTPRequestHandler):
    def do_POST(self):
        body = self.rfile.read(int(self.headers.get("content-length", 0)))
        h = self.headers
        if not verify(body, h.get("x-sume-webhook-timestamp", ""), h.get("x-sume-webhook-signature", "")):
            return self.send_response(401) or self.end_headers()
        if BUSY:
            self.send_response(503); self.send_header("Retry-After", "120")
        else:
            self.send_response(204)  # store body durably first, then answer
        self.end_headers()

HTTPServer(("127.0.0.1", 8080), H).serve_forever()

What are the limits of this trick?

Retry-After only lengthens the wait; it does not pause the clock on the attempt count. Each 503 still uses one of the ten attempts, so a long outage can exhaust delivery even with generous headers. When that happens webhook_delivery.status becomes exhausted and the run stays completed. Read the receipt from result_url, fix the endpoint, and call POST /v1/format-runs/{run_id}/webhook/redeliver.

Because retries repeat the same request_id, dedupe on it and order deliveries by created_at. Sending a 4xx to reject a bad signature is also a failed attempt, so a verifier bug burns the whole budget on valid events.

How should I test the shedding path before production?

Use Send test from /dashboard/webhooks, or POST /v1/webhooks/test-deliveries with account:write, to prove the endpoint is reachable and the signature check passes. The test event is a dummy webhook.test body with no job_id or run id, so your handler must not assume those fields exist. Return 204 for event types you do not recognize rather than a 500, which would turn a harmless new event into a retry storm.

To rehearse the 503 branch, start the sample with BUSY=1, replay a real run with the redeliver route, and confirm the response you get back reports a non-2xx status_code. Then restart without BUSY and redeliver again, which does not use up an automatic attempt.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume