One Sume webhook endpoint for three run types: verify, route
Format, action and agent runs share one signature scheme. Verify HMAC-SHA256 over timestamp.body, refuse an empty secret, then route on the event field.

One endpoint is enough
Sume run webhooks use one signing scheme for three kinds of run, so one endpoint can receive all of them. The events are format.run.terminal, action.run.terminal and agent.run.terminal. Verify the signature first, then read event and send the body to the right handler.
The signature header looks like sume-v1=<hex>. The value is HMAC-SHA256 over the timestamp, a dot, and the raw request body, using your signing secret. Sume rejects a timestamp outside a 5-minute window on your side, so you should too.
A verifier and router
This Python sketch runs with the standard library only. It refuses an empty secret, checks the 300-second window, compares in constant time and returns the handler name. Pass the raw bytes, not a re-serialized dict.
import hashlib, hmac, json, time
HANDLERS = {"format.run.terminal": "format", "action.run.terminal": "action", "agent.run.terminal": "agent"}
def verify(secret, ts, raw, header, now=None):
if not secret:
raise ValueError("empty secret")
now = time.time() if now is None else now
if abs(now - int(ts)) > 300:
return False
mac = hmac.new(secret.encode(), ts.encode() + b"." + raw, hashlib.sha256).hexdigest()
return hmac.compare_digest("sume-v1=" + mac, header)
def route(raw):
return HANDLERS.get(json.loads(raw).get("event"), "ignore")
if __name__ == "__main__":
raw = b'{"event":"format.run.terminal"}'
ts = str(int(time.time()))
sig = "sume-v1=" + hmac.new(b"s3", ts.encode() + b"." + raw, hashlib.sha256).hexdigest()
print(verify("s3", ts, raw, sig), route(raw))What the delivery does
Sume tries up to 10 times, with a 10-second timeout per attempt and no redirects. The delay between attempts is the larger of 30 seconds doubled each time and any Retry-After, capped at one hour. Return a 2xx as soon as the signature checks out, then do the slow work off the request.
| Rule | Value |
|---|---|
| Events | format.run.terminal, action.run.terminal, agent.run.terminal |
| Signature | sume-v1=hex, HMAC-SHA256 of ts.raw_body |
| Attempts | Up to 10 |
| Timeout per attempt | 10 seconds |
| Outcome field | ok, degraded or error |
| No webhook for | canceled and skipped runs |
Edge cases to handle
A receipt over 1 MiB arrives with the payload null and error.result_url, so fetch the result from there. Canceled and skipped runs never send a webhook, so keep a polling sweep for runs that stay silent. If you missed one, POST /v1/format-runs/{id}/webhook/redeliver sends it again.
Treat outcome: degraded as a result to inspect, not a pass. Make your handler idempotent on the run id, since retries can deliver the same event twice.
Testing it
Run the sample file directly. It signs a tiny body with a throwaway secret, verifies it and prints True format. Then test the failures: an old timestamp, a flipped byte in the body and an empty secret should each fail, the last by raising an error. Use the same code in your unit tests so a refactor cannot loosen the check.
Read the body once as bytes, before any JSON parser touches it. Frameworks that re-encode JSON change spacing and break the HMAC.
- Return 2xx fast, work later.
- Dedupe on run id.
- Keep the secret in your secret store, never in the repo.
Make it routine
Add this check to your runbook and run it on a schedule, not only after an incident. The cost is a few read requests, and the rate limits are far above what it needs: even the Free plan allows 120 writes and 4,800 reads per minute. Keep the output with the date, so you can show later what the system looked like when a question came up.
Keep the router small. Map each event to one function, and have every function read the run id, the outcome and, when present, the output. Unknown events should be logged and acknowledged with 2xx, so a new event type added later does not trigger ten retries against a handler that cannot parse it.
- Log the event name and run id.
- Do not log the raw body if it holds customer data.
Sources
Related posts
More in Agents
- Pro key: 300 writes and 12,000 reads a minute for an agent on Sume MCP
On Sume Pro, each key gets 300 writes and 12,000 reads per minute. An agent polling 20 jobs every 5 s uses 240 reads, 2 percent of the read budget.
- Scheduled run spend cap: a per-call number can only lower it
A Sume schedule's default generation cap is $1.00. A per-run number is clamped to the lower of the request and the schedule cap, and 0 is rejected.
- A Sume image job at 50 seconds: slow, queued, or stuck?
Sume's docs say an image still running after a 50 s jobs_wait is usually stuck, not slow. Check queued vs processing and events first; never submit twice.
- Sume MCP avatar tools: 5 read-only and the paid create set
Hosted Sume MCP lists five read tools for avatars and avatar videos and a paid group for creates and previews. What OAuth mcp:read sees and the dry_run flow.
Written by Sume