A 12-second webhook handler times out at 10 s: dedupe Sume deliveries

Sume gives each webhook attempt 10 seconds and retries up to 10 times. A 12 s handler gets repeats. Verify, store job_id once, and return 2xx early.

5 min readSume
All posts

If your webhook handler needs 12 seconds, Sume treats every delivery as failed, because each attempt has a 10-second timeout, and it sends the same terminal event again, up to 10 attempts in all. The fix has two parts: acknowledge with a 2xx only after you have stored the event durably, and use job_id as the idempotency key so a repeat changes nothing. The sample later in this post verifies the signature, stores the job id once, and returns 200 for the repeat.

What Sume does with a slow endpoint

The webhooks guide states that a slow endpoint uses up the attempt budget and Sume retries it. The retry spacing is a fixed delay, 30 seconds by default, not exponential backoff. For a handler that always takes 12 seconds, the table below gives the worst case, with each attempt cut off at 10 seconds on Sume's side. Your server may keep working after Sume has given up on the connection, which is how a single event can start the same expensive work several times in parallel.

Delivery behavior from the Sume webhooks guide, as of 2026-10-08
ItemValueWorst case for a 12 s handler
AttemptsUp to 10 in total10 deliveries of one event
Timeout per attempt10 s10 x 10 s = 100 s of waiting by Sume
SpacingFixed, 30 s by default9 gaps x 30 s = 270 s
Span of the automatic windowAttempts plus gaps100 s + 270 s = 370 s, about 6 min 10 s
Manual replayPOST /v1/jobs/{job_id}/webhook/redeliverDoes not use one of the automatic 10

Verify, store, then return 2xx

Verification follows the docs: sign the string timestamp, a dot, and the raw body with HMAC SHA 256, and compare the result with the sume-v1= entries in x-sume-webhook-signature. During a secret rotation the header carries several entries, so accept the delivery if any one matches. Reject a timestamp outside your tolerance; the docs suggest five minutes. A verifier must also refuse an empty secret, because a missing environment variable would otherwise make every signature predictable. The demo signs its own event, so it runs without a Sume account.

import asyncio, hashlib, hmac, json, sqlite3, time
def verify(secret, ts, body, header, tolerance=300):
    if not secret:
        raise ValueError("empty signing secret")
    if abs(time.time() - int(ts)) > tolerance:
        return False
    digest = hmac.new(secret.encode(), f"{ts}.{body}".encode(), hashlib.sha256).hexdigest()
    return any(hmac.compare_digest(e.strip(), "sume-v1=" + digest)
               for e in header.split(","))

db = sqlite3.connect(":memory:")
db.execute("create table seen (job_id text primary key)")

def handle(secret, ts, body, header):
    if not verify(secret, ts, body, header):
        return 401
    job_id = json.loads(body)["job_id"]
    db.execute("insert or ignore into seen values (?)", (job_id,))
    db.commit()
    return 200

async def main():
    secret, ts = "demo-secret", str(int(time.time()))
    body = json.dumps({"event": "job.completed", "job_id": "job_demo"})
    sig = hmac.new(secret.encode(), f"{ts}.{body}".encode(), hashlib.sha256).hexdigest()
    for attempt in (1, 2):  # the second call is a retry of the same event
        print(attempt, handle(secret, ts, body, "sume-v1=" + sig))
    print(db.execute("select count(*) from seen").fetchone()[0], "stored")

asyncio.run(main())

Move the slow work out of the request

The handler above does only the fast part: verify, insert, return. Do the 12 seconds of work, such as downloading the video and copying it to your own storage, in a separate worker that reads the stored job ids. The primary key on job_id makes the insert a no-op for repeats, so the worker runs once for each job even when Sume delivered the event three times. The same rule covers the event types that a video job can send: job.completed, job.failed, and job.canceled all carry a job_id, and each one is terminal.

  • Return a non-2xx code only for a bad signature or a failed durable write. Sume retries both, and a retry is what you want in those two cases.
  • Keep polling GET /v1/videos/{jobId} as a backup. After ten refused attempts the job is still terminal, but you get no more automatic deliveries.
  • Use redeliver for a job whose attempts ran out, then dedupe on job_id again. The redelivered event carries a fresh timestamp and signature.

Webhook or polling for this job

Use the callback when you want a notification at completion and polling when you want a simple client. They are not exclusive: the jobs guide says to keep the poll in place with webhook mode. On /v1/videos the field is callback_url, it must be a public HTTPS URL, and the payload is Sume's job envelope, not the OpenRouter video.generation envelope.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume