Poison Sume webhook events: park them in a table, return 2xx
A webhook your handler can never process will be retried up to 10 times. Store the raw body in a dead-letter table, return 2xx, and replay it later.

When a Sume webhook arrives that your handler can never process, store the raw body in a dead-letter table, return a 2xx, and deal with it from there. Returning 500 does not make the event processable; it just makes Sume send it again, up to 10 attempts, while the same bad event ties up your receiver. The docs say to return any 2xx after durably storing the event, and a parked row is durable storage.
Poison versus outage
Two different failures look alike from the outside. An outage (your database is down, a deploy is mid-flight) is something retries can fix, and returning a non-2xx is the right answer because the event will arrive again when you are back. A poison event is one where the same input fails the same way every time: a payload you cannot parse, a job_id that your database does not know, a new event name your code predates.
Sume's job webhooks retry on network errors and non-2xx responses until attempts are exhausted, with a 10-second timeout per attempt. For run webhooks, a 1 MiB limit on receipts means large payloads arrive as payload: null plus an error.code of payload_too_large and a result URL, which is another input your handler should park or fetch rather than crash on.
A small dead-letter inbox
The example uses SQLite so it runs anywhere. Every delivery is written to an inbox keyed by job_id with insert or ignore, so a repeat delivery or a manual redelivery is a no-op. Anything that cannot be parsed is parked with a reason, and the function returns 204 in both cases. A worker later reads status = 'new' rows to do the real work, and a person reads the parked ones.
Verify the signature before you store anything, on the raw body; unsigned or mismatched requests deserve a 401, not a parking spot.
import json, sqlite3
db = sqlite3.connect(":memory:")
db.execute("create table inbox (id text primary key, body text, status text, reason text)")
def receive(raw: bytes) -> int:
try:
event = json.loads(raw)
key = event["job_id"]
except (ValueError, KeyError):
db.execute("insert into inbox values (?, ?, 'parked', 'unparseable')",
(f"bad-{db.execute('select count(*) from inbox').fetchone()[0]}", raw.decode(errors="replace")))
return 204 # stored, so stop the retries
db.execute("insert or ignore into inbox values (?, ?, 'new', null)", (key, raw.decode()))
return 204
print(receive(b'{"event":"job.completed","job_id":"job_1"}'))
print(receive(b'{"event":"job.completed","job_id":"job_1"}'))
print(receive(b"not json"))
print(db.execute("select id, status, reason from inbox").fetchall())Replaying parked events
A parked event is not lost, because the job or run still reached its true terminal state. When you fix your handler, you can reprocess the stored body directly, or ask Sume to send it again: POST /v1/jobs/{job_id}/webhook/redeliver with a jobs:write key re-sends the real terminal event with a fresh timestamp and does not consume one of the automatic 10. Because your inbox is keyed by job_id, the redelivery lands on the existing row instead of making a second.
| Situation | Why | What to return |
|---|---|---|
| Database down during deploy | Retry fixes it | Return 5xx or 503 |
| Body is not valid JSON | Never parseable | Park, return 2xx |
| Unknown event name | Your code is older | Return 204 and log |
| Unknown job_id | May be another environment | Park, return 2xx, investigate |
| Payload null, payload_too_large | Known case | Fetch the result URL, then store |
Alert on the table, not the endpoint
Count parked rows per hour and alert when the number is above zero for more than a day. That is a more useful signal than endpoint 5xx rates, which will stay quiet because you now answer 2xx. Keep the raw body, the event name and the delivery timestamp header; you will want them when you reproduce the failure.
What not to park
Do not park events that failed signature verification: those are not Sume's deliveries, and storing them lets anyone fill your table. Do not park events your own database rejected for being temporarily unavailable either; those should return a non-2xx so Sume retries. Park only what you have decided is permanently unprocessable, and give each parked row a reason string you can group by. A small weekly review of the reasons is usually enough to find the bug in your handler.
Sources
Related posts
More in Developers
- Per-customer spend caps on Sume: what Sume caps, what you log
A Sume key belongs to a workspace, not your end customer. Cap each run with generation_spend_cap_usd and keep a per-customer ledger yourself.
- Perplexity Decisions API as a publish gate for Sume output
Check a finished Sume Format image with Perplexity's Decisions API before it ships: base64 data URL, one yes/no question, a threshold, and a human-review lane.
- Pick the cheapest Sume image model for an aspect ratio (Python)
A short Python script that reads GET /v1/images/models, keeps models that list your ratio, reference count and n, then prices them from the endpoint records.
- Pick a video model from GET /v1/videos/models, not a constant
A new video model should not need a redeploy. Filter the Sume catalog by resolution, duration, audio and frame inputs at runtime, and survive a withdrawn model.
Written by Sume