Load-test your Sume webhook receiver with a signed burst

Fire 500 correctly signed job.completed requests at your own receiver before a bulk run does it for real. A runnable Python harness and the 10-second budget.

6 min readSume
All posts

Before you launch a 100-item bulk run, replay a burst of correctly signed job.completed requests at your own receiver and check two numbers: every response is a 2xx, and the slowest one is far below 10 seconds. Sume allows 10 seconds per delivery attempt, and a slow endpoint burns the attempt budget and gets retried, so a receiver that is merely slow under load turns into duplicate deliveries. Doing the test with your own signatures costs no credits and touches no Sume endpoint.

Why a burst is the realistic case

A bulk run holds 1 to 100 items and runs 1 to 16 at a time. Items can finish in clumps, and each item can carry its own communication.webhook_url, because the queue itself has no webhook. So the pattern your receiver sees is not a steady trickle but a wave of up to a few dozen terminal events within seconds.

Each delivery is an HMAC check over <timestamp>.<raw_body>, a database write and a 2xx. None of that is heavy, but the failure modes are: a connection pool of 5, a cold serverless start, a synchronous call to storage before the response. You want to find those on a Tuesday afternoon, not mid-campaign.

A harness you can run

The script starts a throwaway verifier on localhost, signs 500 events with the documented scheme (x-sume-webhook-timestamp, x-sume-webhook-signature: sume-v1=<hex>), sends them with 32 threads and prints the set of status codes and the slowest latency. It refuses an empty secret on startup. Point url at your staging receiver and import your own verifier to test the real thing; the one inside the script is a stand-in so it runs by itself on Python 3.

It uses only the standard library.

import hashlib, hmac, json, threading, time, urllib.request
from concurrent.futures import ThreadPoolExecutor
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

SECRET = "whsec_test_only"
assert SECRET, "refuse an empty secret"

def sign(ts, raw):
    mac = hmac.new(SECRET.encode(), f"{ts}.".encode() + raw, hashlib.sha256)
    return "sume-v1=" + mac.hexdigest()

class H(BaseHTTPRequestHandler):
    def do_POST(self):
        raw = self.rfile.read(int(self.headers["content-length"]))
        ts = self.headers["x-sume-webhook-timestamp"]
        ok = hmac.compare_digest(sign(ts, raw), self.headers["x-sume-webhook-signature"])
        self.send_response(204 if ok else 401)
        self.end_headers()
    def log_message(self, *a): pass

srv = ThreadingHTTPServer(("127.0.0.1", 0), H)
threading.Thread(target=srv.serve_forever, daemon=True).start()
url = f"http://127.0.0.1:{srv.server_port}/hook"

def fire(i):
    raw = json.dumps({"event": "job.completed", "job_id": f"job_{i}"}).encode()
    ts = str(int(time.time()))
    req = urllib.request.Request(url, raw, {"x-sume-webhook-timestamp": ts, "x-sume-webhook-signature": sign(ts, raw)})
    t0 = time.perf_counter()
    code = urllib.request.urlopen(req, timeout=10).status
    return code, time.perf_counter() - t0

with ThreadPoolExecutor(32) as pool:
    out = list(pool.map(fire, range(500)))
print(sorted({c for c, _ in out}), "max seconds", round(max(t for _, t in out), 3))

How to read the result

The status set should be [204] alone. A 401 among the codes means your verifier disagrees with the signer, usually because something parsed and re-serialised the JSON before hashing. Hash the raw bytes, as the webhooks page says. A 5xx or a timeout means the receiver cannot sustain the burst, and each of those would have been retried by Sume, up to 10 attempts.

read 2026-10-03
ObservationMeaningNext step
Status set is [204]Receiver keeps upMove on to idempotency checks
Any 401Signature mismatchHash raw bytes, check secret and tolerance
Any 5xx or timeoutReceiver overloadedAcknowledge first, work in a queue
Max latency above 2 sToo close to the 10 s budgetProfile the write path

What the test cannot tell you

It does not exercise rotation windows, where the signature header carries two entries, newest first, so add a case with sume-v1=<new>,sume-v1=<old>. It also does not test replay: send the same job_id twice and assert only one row is written, since deliveries and manual redelivers can both repeat. Finally, run it at the receiver you will actually deploy, with the same proxy in front of it.

Making the burst realistic

Vary the payloads: include job.failed with an error object, a job.canceled, a run event with an outcome of degraded, and a receipt with payload: null for the over-1-MiB case. Send a few duplicates and a few deliveries out of order, and assert the stored state is the same whichever order they arrive in, ordering by created_at where you need an order. Finally, run the burst while a real deploy is rolling, which is when the 10-second timeout matters most.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume