Webhook for an unknown job_id: park it, then reconcile on insert
A Sume job.completed webhook can name a job_id your database has not stored yet. Return 2xx, park the event in SQLite, and apply it when the insert lands.

If a verified Sume webhook names a job_id that is not in your database, do not return 404 and do not drop it. Store the event in a parking table, return 2xx, and apply it when the job row is inserted. Sume's delivery guide says to return 2xx only after a durable store, and a parked row is a durable store. A 5xx or a timeout only makes Sume retry the same event on its schedule, which is 10 attempts with a fixed 30 second spacing per Sume webhooks.
The case is rarer than it sounds for video, which takes minutes, but it is real in three situations: your submit worker crashed after the POST returned and before the INSERT, a read replica lags the primary, or a very fast job finishes while your transaction is still open. In each case the fix is the same small pattern, and it costs two tables and about twenty lines of code.
What does the handler need to store?
The payload carries event (job.completed, job.failed or job.canceled), request_id, job_id, a status of OK or ERROR, and payload.artifacts on success. Use job_id as the idempotency key for the whole handler, as the webhook docs recommend. That means two tables: jobs for rows you created and parked for events that arrived first.
Parking is an upsert on job_id, so a second delivery of the same event overwrites nothing. Draining is the reverse step: in the same transaction that inserts the job, look for a parked row, apply it, and delete it.
Can I see it run end to end?
import sqlite3
db = sqlite3.connect(":memory:")
db.executescript("""
create table jobs(job_id text primary key, status text, url text);
create table parked(job_id text primary key, event text, url text);""")
def on_event(job_id, event, url=None):
if db.execute("select 1 from jobs where job_id=?", (job_id,)).fetchone():
st = "completed" if event == "job.completed" else "failed"
db.execute("update jobs set status=?, url=? where job_id=?", (st, url, job_id))
outcome = "applied"
else:
db.execute("insert or ignore into parked values (?,?,?)", (job_id, event, url))
outcome = "parked"
db.commit()
return outcome
def register(job_id):
with db:
db.execute("insert or ignore into jobs values (?, 'queued', null)", (job_id,))
q = "select event, url from parked where job_id=?"
row = db.execute(q, (job_id,)).fetchone()
if row:
on_event(job_id, *row)
db.execute("delete from parked where job_id=?", (job_id,))
assert on_event("job_1", "job.completed", "https://cdn.example/a.mp4") == "parked"
register("job_1")
print(db.execute("select * from jobs").fetchall())
assert db.execute("select status from jobs").fetchone()[0] == "completed"How do I keep orphans from piling up?
Watch the parked table. A row that sits there for more than a few minutes means the job was created on Sume but never recorded by you, which is an orphan that may be billing. Alert on the age of the oldest parked row, and when one is found, look the job up by job_id with GET /v1/jobs/{id} to see what it was before you decide whether to adopt it or leave it. Because a submit retried with the same Idempotency-Key returns the original job, a crash-and-retry usually heals itself and the parked event drains on the second insert.
Do not generate your own job ids to avoid the race. The id is Sume's request_id, returned in the 202 envelope, and it is what the webhook carries back. Your own correlation id belongs in your table next to it, not in place of it.
What if the event never comes?
Parking is only half of the repair. If the event never arrives, the job row you inserted sits at queued forever. Add a sweeper that reads GET /v1/jobs/{id}/status for any row older than your expected render time and writes whatever terminal state it returns. Sume documents polling and webhooks as two views of the same job, so the sweeper and the handler can both write the same final state safely when the write is conditional on the row not already being terminal.
Unknown events are a different case. An event name you do not know, such as a future type, should return 204, not 500, so that Sume does not retry something you cannot act on. An unknown job id for a known event is a race, so you park it. Keep both rules in the same router, covered by one test each, and keep signature verification in front of both: an unsigned request should never reach the parking table, because that table would then be writable by anyone.
| Case | Response | Why |
|---|---|---|
| Bad or missing signature | 401 | Nothing is stored |
| Known event, job row found | 200 after the update | Durable |
| Known event, job row missing | 200 after parking | Durable, reconciled later |
| Unknown event name | 204 | Not retried for nothing |
Sources
Related posts
More in Developers
- Sume TTS source errors: which are safe to retry (status table)
Every tts_ error code in Sume's script-source API with its HTTP status, whether it charges, and whether a retry can help. A table for client error handling.
- Word timestamps to video frame numbers at 29.97 fps in Python
Sume STT words[] carry start and end in seconds. Convert them to frame indexes with exact 30000/1001 math so cuts do not drift on long timelines.
- Zed context_servers for the hosted Sume server: no header means OAuth
Add the hosted Sume server to Zed's settings.json context_servers with a url. With no Authorization header Zed runs the MCP OAuth flow, so start with mcp:read.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
Written by Sume