Can a Sume webhook arrive twice? Build an idempotent receiver

Sume retries failed webhook deliveries up to 10 times and Redeliver replays a real event, so one terminal event can reach you twice. Dedupe on job_id or run_id.

5 min readSume
All posts

Yes, plan for it. Sume sends one terminal event for a job or a run, but the delivery path can send the same body more than once: any non-2xx answer or timeout is retried, and the Redeliver action replays the real event on request. Your receiver must therefore treat job_id (generation jobs) or request_id (runs) as an idempotency key and make the second arrival a no-op.

This page shows which fields to dedupe on, a receiver you can run, and the order of operations that keeps a crash from losing an event. The signature check is the same one the Webhooks page documents, so one verifier covers both surfaces.

Where the second copy comes from

The docs name three paths that put an event on your server more than once. None of them is a bug, and none of them changes the job: a delivery failure leaves the job in its real terminal state.

  • Automatic retries. Sume retries network errors and non-2xx responses, up to 10 attempts in total, with a 10-second timeout on each attempt. If your endpoint stored the event but answered slowly, the retry still comes.
  • Manual Redeliver. POST /v1/jobs/{job_id}/webhook/redeliver (needs jobs:write) re-POSTs the real terminal event with a fresh timestamp and signature. For Format runs the route is POST /v1/format-runs/{run_id}/webhook/redeliver (needs formats:write). It still works after the automatic attempts are exhausted.
  • Your own replays. A queue consumer that crashes after storing but before acking will see its own copy again. That one is yours, not Sume's, but the same dedupe key fixes it.

Which field is the dedupe key

Job webhooks and run webhooks have different payloads but the same signature scheme. The key differs by surface.

Dedupe keys by webhook surface (read 2026-10-07)
SurfaceEventsDedupe onOrder deliveries by
Generation jobjob.completed, job.failed, job.canceledjob_idNot documented. Read the job from status_url for the true state.
Format runformat.run.terminalrequest_id (equals run_id, stable across retries)created_at (build time of that delivery body)
Action runaction.run.terminalrequest_idcreated_at
Agent Completionagent.run.terminalrequest_idcreated_at

A receiver that survives duplicates

The receiver below verifies the sume-v1 signature over <timestamp>.<raw_body>, refuses an empty secret, accepts any of the comma-separated signatures that appear during a secret rotation, and then inserts the key with insert or ignore. A duplicate returns a different 2xx, which is enough to stop Sume retrying. Swap SQLite for your own table; the point is a unique key on the event id.

import hashlib, hmac, json, sqlite3, time

db = sqlite3.connect(":memory:")
db.execute("create table seen (id text primary key, body text, done int default 0)")

def verify(raw: bytes, ts: str, header: str, secret: str, tolerance: int = 300) -> bool:
    if not secret or not ts.isdigit() or abs(time.time() - int(ts)) > tolerance:
        return False
    mac = hmac.new(secret.encode(), ts.encode() + b"." + raw, hashlib.sha256).hexdigest()
    return any(hmac.compare_digest(p.strip(), "sume-v1=" + mac) for p in header.split(","))

def receive(raw: bytes, headers: dict, secret: str) -> int:
    if not verify(raw, headers.get("x-sume-webhook-timestamp", ""),
                  headers.get("x-sume-webhook-signature", ""), secret):
        return 401
    event = json.loads(raw)
    key = event.get("job_id") or event["request_id"]
    cur = db.execute("insert or ignore into seen (id, body) values (?, ?)", (key, raw.decode()))
    db.commit()
    return 204 if cur.rowcount == 0 else 202  # both are 2xx: Sume stops retrying

if __name__ == "__main__":
    body = json.dumps({"event": "job.completed", "job_id": "job_1", "request_id": "job_1"}).encode()
    ts = str(int(time.time()))
    sig = "sume-v1=" + hmac.new(b"s3cret", ts.encode() + b"." + body, hashlib.sha256).hexdigest()
    h = {"x-sume-webhook-timestamp": ts, "x-sume-webhook-signature": sig}
    print(receive(body, h, "s3cret"), receive(body, h, "s3cret"), receive(body, h, ""))

Store first, then answer, then work

Sume counts any 2xx as delivered and gives each attempt 10 seconds. The safe order is: verify, write the raw event to durable storage, return 2xx, and only then do the slow work (copying the MP4, calling your CRM, emailing a customer).

If you return 2xx before the write is durable, a crash loses the event and Sume will not resend it. If you do the slow work before answering, a slow handler uses the 10-second budget and Sume retries a delivery you were already processing. Add a done column so the worker that picks up the stored row can tell a finished event from an unfinished one; a Redeliver of an unfinished row then simply re-queues the work.

What a duplicate must not do

Idempotent means the second arrival changes nothing, including your billing and notifications. Do not send a second email, do not append a second row to a customer-facing list, and do not start a second downstream job. If the first delivery failed halfway through your own processing, the stored row with done = 0 is how you resume, not a reason to process the body twice in parallel.

Keep the poll path as the backup. A webhook is a delivery optimization, and the docs say to keep status_url polling available for events that never arrive. A nightly sweep that reads each non-terminal job and feeds it through the same insert or ignore path closes the gap without a second code path.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume