Load-test your Sume webhook receiver with a signed burst
Fire 500 correctly signed job.completed requests at your own receiver before a bulk run does it for real. A runnable Python harness and the 10-second budget.

Before you launch a 100-item bulk run, replay a burst of correctly signed job.completed requests at your own receiver and check two numbers: every response is a 2xx, and the slowest one is far below 10 seconds. Sume allows 10 seconds per delivery attempt, and a slow endpoint burns the attempt budget and gets retried, so a receiver that is merely slow under load turns into duplicate deliveries. Doing the test with your own signatures costs no credits and touches no Sume endpoint.
Why a burst is the realistic case
A bulk run holds 1 to 100 items and runs 1 to 16 at a time. Items can finish in clumps, and each item can carry its own communication.webhook_url, because the queue itself has no webhook. So the pattern your receiver sees is not a steady trickle but a wave of up to a few dozen terminal events within seconds.
Each delivery is an HMAC check over <timestamp>.<raw_body>, a database write and a 2xx. None of that is heavy, but the failure modes are: a connection pool of 5, a cold serverless start, a synchronous call to storage before the response. You want to find those on a Tuesday afternoon, not mid-campaign.
A harness you can run
The script starts a throwaway verifier on localhost, signs 500 events with the documented scheme (x-sume-webhook-timestamp, x-sume-webhook-signature: sume-v1=<hex>), sends them with 32 threads and prints the set of status codes and the slowest latency. It refuses an empty secret on startup. Point url at your staging receiver and import your own verifier to test the real thing; the one inside the script is a stand-in so it runs by itself on Python 3.
It uses only the standard library.
import hashlib, hmac, json, threading, time, urllib.request
from concurrent.futures import ThreadPoolExecutor
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
SECRET = "whsec_test_only"
assert SECRET, "refuse an empty secret"
def sign(ts, raw):
mac = hmac.new(SECRET.encode(), f"{ts}.".encode() + raw, hashlib.sha256)
return "sume-v1=" + mac.hexdigest()
class H(BaseHTTPRequestHandler):
def do_POST(self):
raw = self.rfile.read(int(self.headers["content-length"]))
ts = self.headers["x-sume-webhook-timestamp"]
ok = hmac.compare_digest(sign(ts, raw), self.headers["x-sume-webhook-signature"])
self.send_response(204 if ok else 401)
self.end_headers()
def log_message(self, *a): pass
srv = ThreadingHTTPServer(("127.0.0.1", 0), H)
threading.Thread(target=srv.serve_forever, daemon=True).start()
url = f"http://127.0.0.1:{srv.server_port}/hook"
def fire(i):
raw = json.dumps({"event": "job.completed", "job_id": f"job_{i}"}).encode()
ts = str(int(time.time()))
req = urllib.request.Request(url, raw, {"x-sume-webhook-timestamp": ts, "x-sume-webhook-signature": sign(ts, raw)})
t0 = time.perf_counter()
code = urllib.request.urlopen(req, timeout=10).status
return code, time.perf_counter() - t0
with ThreadPoolExecutor(32) as pool:
out = list(pool.map(fire, range(500)))
print(sorted({c for c, _ in out}), "max seconds", round(max(t for _, t in out), 3))How to read the result
The status set should be [204] alone. A 401 among the codes means your verifier disagrees with the signer, usually because something parsed and re-serialised the JSON before hashing. Hash the raw bytes, as the webhooks page says. A 5xx or a timeout means the receiver cannot sustain the burst, and each of those would have been retried by Sume, up to 10 attempts.
| Observation | Meaning | Next step |
|---|---|---|
| Status set is [204] | Receiver keeps up | Move on to idempotency checks |
| Any 401 | Signature mismatch | Hash raw bytes, check secret and tolerance |
| Any 5xx or timeout | Receiver overloaded | Acknowledge first, work in a queue |
| Max latency above 2 s | Too close to the 10 s budget | Profile the write path |
What the test cannot tell you
It does not exercise rotation windows, where the signature header carries two entries, newest first, so add a case with sume-v1=<new>,sume-v1=<old>. It also does not test replay: send the same job_id twice and assert only one row is written, since deliveries and manual redelivers can both repeat. Finally, run it at the receiver you will actually deploy, with the same proxy in front of it.
Making the burst realistic
Vary the payloads: include job.failed with an error object, a job.canceled, a run event with an outcome of degraded, and a receipt with payload: null for the over-1-MiB case. Send a few duplicates and a few deliveries out of order, and assert the stored state is the same whichever order they arrive in, ordering by created_at where you need an order. Finally, run the burst while a real deploy is rolling, which is when the 10-second timeout matters most.
Sources
Related posts
More in Developers
- LRC lyrics file from an AI music track: Sume STT segments in Python
YouTube accepts .lrc lyric files. A short Python script turns STT sentence segments from a generated song into [mm:ss.xx] lines. Check results before upload.
- Luma Agents API polling: 20 s wait, 10-minute video timeout, and Sume
Luma's quickstart says wait 20 seconds, then poll, with a hard 2-minute image and 10-minute video timeout, not backoff. A matching Sume loop that actually runs.
- Luma Ray 2 to Ray 3.2 migration: the field map and what changes
Luma is retiring Ray 3, Ray 2, Ray 2 Flash and more for ray-3.2 on one /v1/generations endpoint. The field map, the poll-only change, and Sume's contrast.
- macos-14 brownouts from Oct 5: rerun a Sume submit with the same key
GitHub's macos-14 runners fail on purpose in brownout windows before the 2026-11-02 retirement. Derive a stable Idempotency-Key so a re-run does not bill twice.
Written by Sume