Webhook handler slower than 10 seconds: duplicate Sume deliveries

Sume gives each webhook attempt 10 seconds, then retries up to 10 times, 30s apart. A slow handler that succeeds late still gets duplicates. Ack first.

5 min readSume
All posts

Sume waits 10 seconds for each webhook attempt. If your handler is still working at that point, the attempt counts as failed and Sume retries, so a handler that finishes in 14 seconds can receive the same job.completed more than once. Return a 2xx right after durably storing the event, and do the slow work afterwards.

This follows the delivery rules on the Webhooks page: up to 10 attempts in total, a fixed delay between attempts (30s by default, not exponential backoff), and a 10s timeout per attempt.

What exactly happens when the handler is slow?

The docs state that a slow endpoint burns the attempt budget and gets retried. From Sume's side a timed-out attempt looks the same as a refused one, even if your code later commits successfully.

Ten failed attempts leave a failed delivery, while the job still reached its real terminal state. Delivery is an optimization, and status_url polling stays available for events that never arrive.

Job webhook delivery limits (docs read 2026-10-02)
SettingValue
AttemptsUp to 10 in total
SpacingFixed, 30s by default
Per-attempt timeout10s
SuccessAny 2xx after you store the event
Idempotency key on your sidejob_id

How do I make duplicates harmless?

The docs tell receivers to use job_id as the idempotency key. For run webhooks the envelope request_id equals the run id and is stable across retries, so dedupe on that and use created_at if you need ordering.

The pattern: verify the signature on the raw body, insert the event keyed by id (ignore a conflict), return 204, and let a worker do the rest.

Ack-then-work receiver

The sketch uses only the standard library, verifies the HMAC over <timestamp>.<raw_body>, and refuses to run with an empty secret. Accept the delivery if any sume-v1= entry matches, since the header carries one entry per live secret during a rotation.

import hashlib, hmac, os, queue, time

SECRET = os.environ.get("SUME_COM_WEBHOOK_SIGNING_SECRET", "")
if not SECRET:
    raise SystemExit("signing secret is empty")
WORK: queue.Queue = queue.Queue()
SEEN: set = set()

def handle(raw: bytes, ts: str, sig_header: str) -> int:
    if abs(time.time() - int(ts)) > 300:
        return 401
    mac = hmac.new(SECRET.encode(), ts.encode() + b"." + raw, hashlib.sha256)
    want = "sume-v1=" + mac.hexdigest()
    ok = False
    for part in sig_header.split(","):
        ok |= hmac.compare_digest(part.strip(), want)
    if not ok:
        return 401
    import json
    event = json.loads(raw)
    key = event.get("job_id") or event.get("request_id")
    if key not in SEEN:
        SEEN.add(key)
        WORK.put(event)
    return 204

Can I trigger a missed delivery again?

Yes. POST /v1/jobs/{job_id}/webhook/redeliver (scope jobs:write) re-POSTs the real terminal event with a fresh timestamp and signature, and it does not consume one of the automatic 10. Redeliver does not change the destination URL. See Run webhooks for the run equivalent.

What Sume does not do

Sume does not follow redirects on run webhook URLs, and it does not send progress events for jobs. Use the SEEN set above only as a demo: in production store keys in a database so a restart does not forget them.

Quick checklist

The points above reduce to a short list you can paste into a runbook.

  • Verify the signature on the raw body before parsing JSON.
  • Reject timestamps outside your replay window; five minutes is the documented default.
  • Store the event, return 2xx, and do the work afterwards.
  • Dedupe on job_id for jobs and request_id for runs.
  • Keep status_url polling available for events that never arrive.
  • Use redeliver for a missed terminal event; it does not use up the automatic attempts.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume