Dedupe Sume webhooks with SQLite: INSERT OR IGNORE on job_id and event

Sume retries up to 10 times and Redeliver repeats events. A stdlib SQLite table keyed on job_id and event makes your video webhook handler safe to run twice.

4 min readSume
All posts

Make a Sume webhook handler safe to run twice by writing each delivery into a table whose primary key is job_id plus event, using INSERT OR IGNORE, and doing the real work only when the insert changed a row. Sume retries failed deliveries up to ten times and its Redeliver action re-posts terminal events, so the same job_id will arrive again, and that is normal.

Why duplicates are guaranteed

Sume sends terminal job events only: job.completed, job.failed and job.canceled. Each attempt has a 10-second timeout, and any non-2xx response or timeout is retried until the 10-attempt budget is gone, with a fixed spacing of 30 seconds by default. If your server stored the event but timed out before it answered, the next attempt carries the same event.

Redeliver is a second source. It re-posts the real terminal event of a job, with a fresh timestamp and signature, and it does not use one of the automatic attempts. The Sume docs say receivers must treat job_id as the idempotency key. A fresh timestamp means a replay-window check will not save you, and only your own table will.

What to key on

Key on the pair of job id and event name. A job emits one terminal event in practice, but the pair makes the rule explicit and protects you if you also store other event types later. OpenRouter's video guide uses a similar idea, an idempotency header built from the job id and status, but Sume's envelope has no such header, so build it yourself from the body.

Duplicate sources for a Sume job webhook (Sume docs, read 2026-10-05)
SourceSame job_id againTimestampDedupe by
Retry after non-2xx or timeoutYesNew for each attemptjob_id and event
Redeliver actionYesFreshjob_id and event
Secret rotation overlapNo, one deliverySameAccept any signature
Send testNo job_idFreshIgnore webhook.test

The table and the handler

The function returns True only for the first time it sees a delivery. Call it after verification and before any slow work. SQLite's rowcount after INSERT OR IGNORE is 1 for a new row and 0 for a duplicate, which is the whole trick.

import sqlite3

db = sqlite3.connect(":memory:")  # use a file path in production
db.execute("""CREATE TABLE IF NOT EXISTS deliveries (
    job_id TEXT NOT NULL, event TEXT NOT NULL,
    received_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP,
    PRIMARY KEY (job_id, event))""")

def first_time(event: dict) -> bool:
    if event.get("event") == "webhook.test" or not event.get("job_id"):
        return False
    cur = db.execute("INSERT OR IGNORE INTO deliveries (job_id, event) VALUES (?, ?)",
                     (event["job_id"], event["event"]))
    db.commit()
    return cur.rowcount == 1

e = {"event": "job.completed", "job_id": "job_demo_1", "status": "OK"}
print(first_time(e))   # True, do the work
print(first_time(e))   # False, a retry or a redeliver
print(first_time({"event": "job.failed", "job_id": "job_demo_1"}))  # True, different event
print(first_time({"event": "webhook.test", "request_id": "req_wh_test_1"}))  # False

Do the work after the insert, not before

The order is the important part. If you do the work first and insert later, a crash between the two repeats the work. If you insert first and the work crashes, you lose it. The safe design uses the table as an inbox: insert, return 204, and let a worker process rows with a processed flag and a retry count. Then the table is both your dedupe and your queue, and a crash anywhere leaves a row you can pick up.

For a video job the work is usually a download. The artifact URL comes from the poll route, GET /v1/videos/{id}/content, so the worker can fetch it whenever it runs, with your key held on the server side.

Checklist

  • Verify the signature first. Do not insert unauthenticated events.
  • Return 2xx quickly, since Sume only counts a 2xx as delivered.
  • Ignore webhook.test, which has no job_id and never replays a real job.
  • Keep rows for at least as long as a job could be redelivered, then prune.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume