Webhook returns 503 with Retry-After: how Sume spaces the retries
If your webhook endpoint answers 429 or 503 with Retry-After, Sume run webhooks wait at least that long, capped at one hour. Here is how to shed load safely.

When a Sume run webhook reaches an endpoint that answers 429 or 503 with a Retry-After header, Sume waits at least that long before the next attempt, up to a cap of one hour. The documented wait is min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h), so a short Retry-After never makes Sume retry sooner than its own backoff.
That makes a temporary 503 a legitimate way to shed load during a spike, as long as you remember the budget is still ten attempts. Sources: Run webhooks and Runs and results, read 2026-10-02.
Does Retry-After apply to job webhooks too?
The docs I read describe Retry-After handling only for run webhooks. The generation-job page describes a fixed delay between attempts, 30 seconds by default, and does not mention Retry-After. If you need a pause on a job webhook, do not depend on the header; accept the event, store it and process later.
What counts as a failed attempt?
Anything that is not a timely 2xx uses up an attempt.
| Your endpoint | Result |
|---|---|
| 2xx within 10 seconds | Delivered |
| 429 or 503 with Retry-After | Failed attempt; next wait is the longer of the backoff and Retry-After, capped at 1 hour |
| Slow response over 10 seconds | Timed out; retried |
| 3xx redirect | Not followed; counts as a failed attempt |
| Other non-2xx or network error | Failed attempt; retried until ten are used |
How do I shed load without dropping events?
The safest pattern is to verify, write the raw event to durable storage, and answer 204 immediately. Only answer 503 with Retry-After when you cannot even store it. Below is a minimal stdlib receiver that refuses to start without a secret, checks the signature and replay window, and sheds load when told to. During a secret rotation the signature header carries several comma-separated entries, so it accepts any sume-v1= match.
import hashlib, hmac, os, time
from http.server import BaseHTTPRequestHandler, HTTPServer
SECRET = os.environ.get("SUME_COM_WEBHOOK_SIGNING_SECRET", "")
if not SECRET:
raise SystemExit("set SUME_COM_WEBHOOK_SIGNING_SECRET")
BUSY = os.environ.get("BUSY") == "1"
def verify(body, ts, header):
if not ts.isdigit() or abs(time.time() - int(ts)) > 300:
return False
mac = hmac.new(SECRET.encode(), ts.encode() + b"." + body, hashlib.sha256)
want = "sume-v1=" + mac.hexdigest()
return any(hmac.compare_digest(e.strip(), want) for e in header.split(","))
class H(BaseHTTPRequestHandler):
def do_POST(self):
body = self.rfile.read(int(self.headers.get("content-length", 0)))
h = self.headers
if not verify(body, h.get("x-sume-webhook-timestamp", ""), h.get("x-sume-webhook-signature", "")):
return self.send_response(401) or self.end_headers()
if BUSY:
self.send_response(503); self.send_header("Retry-After", "120")
else:
self.send_response(204) # store body durably first, then answer
self.end_headers()
HTTPServer(("127.0.0.1", 8080), H).serve_forever()What are the limits of this trick?
Retry-After only lengthens the wait; it does not pause the clock on the attempt count. Each 503 still uses one of the ten attempts, so a long outage can exhaust delivery even with generous headers. When that happens webhook_delivery.status becomes exhausted and the run stays completed. Read the receipt from result_url, fix the endpoint, and call POST /v1/format-runs/{run_id}/webhook/redeliver.
Because retries repeat the same request_id, dedupe on it and order deliveries by created_at. Sending a 4xx to reject a bad signature is also a failed attempt, so a verifier bug burns the whole budget on valid events.
How should I test the shedding path before production?
Use Send test from /dashboard/webhooks, or POST /v1/webhooks/test-deliveries with account:write, to prove the endpoint is reachable and the signature check passes. The test event is a dummy webhook.test body with no job_id or run id, so your handler must not assume those fields exist. Return 204 for event types you do not recognize rather than a 500, which would turn a harmless new event into a retry storm.
To rehearse the 503 branch, start the sample with BUSY=1, replay a real run with the redeliver route, and confirm the response you get back reports a non-2xx status_code. Then restart without BUSY and redeliver again, which does not use up an automatic attempt.
Sources
Related posts
More in Developers
- Webhook IP allowlist: GitHub and Stripe publish ranges, Sume does not
GitHub exposes webhook IPs via /meta and Stripe lists addresses. The Sume docs publish no sender ranges, so verify the signature and a 5-minute window.
- Return 503 with Retry-After in a deploy: Sume run webhooks honour it
Sume run webhooks honour Retry-After on 429 and 503 up to 1 hour. Job webhooks use a fixed 30s spacing, so keep the drain short and the fallback poll on.
- Sume webhook_url 400: localhost, http, credentials, alias clash
Why a Sume create call rejects webhook_url with 400 invalid_request: localhost, private IPs, http, credentials in the URL, or conflicting webhook_url aliases.
- Which Sume audio endpoint to call: TTS, STT, music, detach, timeline
A decision map for Sume's audio API: seven endpoints, what each takes in and returns, limits and list prices, and the order they chain in.
Written by Sume