Missed Sume video webhook? Poll first, redeliver second

A missed callback does not mean a lost job. A Python sweeper polls open ids, then calls POST /v1/jobs/{id}/webhook/redeliver for the terminal ones.

4 min readSume
All posts

When a Sume webhook never arrives, the job still reached its real terminal state, so reconcile from the job record: poll every id your database still marks as open, and for the ones that are terminal, either handle them directly or ask for a redelivery with POST /v1/jobs/{job_id}/webhook/redeliver. The sweeper below does the first step and prints the redeliver calls for the second.

Webhook is an optimization

Sume says it plainly: a webhook is a delivery optimization, not your only recovery path. After ten refused attempts you have a failed delivery and a job that is still completed, failed or canceled. The job object and the job events carry the webhook delivery state, with an attempt count, when it is available, so you can see which stage broke.

The sweeper is the other half of the design. A webhook gives you speed, and a sweep gives you correctness. Together they give you a pipeline that is fast on a good day and complete on a bad one.

Poll first, redeliver second

Polling is a read and Redeliver is a write that needs the jobs:write scope, so the order matters for both cost and permissions. Often you do not need a redelivery at all, because you can handle the terminal state from the poll response yourself. Use Redeliver when you want the same code path that a live webhook takes.

Recovery options for a missed callback (Sume docs, read 2026-10-05)
OptionRouteScopeUses an automatic attempt
Read the jobGET /v1/videos/{id}readNot applicable
Read the eventsGET /v1/jobs/{id}/eventsreadNot applicable
Redeliver the terminal eventPOST /v1/jobs/{job_id}/webhook/redeliverjobs:writeNo
Send testPOST /v1/webhooks/test-deliveriesaccount:writeNot a real job

The sweeper

Keep your open ids in a dict with a submit time. Anything younger than a grace period is left alone, since the live webhook may still be on its way. Everything else is read from the job route.

import json, os, time, urllib.request, urllib.error
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}

def call(method, path):
    req = urllib.request.Request("https://api.sume.com" + path, method=method, headers=H,
                                 data=b"{}" if method == "POST" else None)
    try:
        with urllib.request.urlopen(req, timeout=30) as r:
            return r.status, json.loads(r.read() or b"{}")
    except urllib.error.HTTPError as e:
        return e.code, json.loads(e.read() or b"{}")

def reconcile(open_jobs: dict, grace_s: int = 600, push: bool = False):
    for job_id, t0 in list(open_jobs.items()):
        if time.time() - t0 < grace_s:
            continue
        code, job = call("GET", f"/v1/videos/{job_id}")
        if code != 200 or job.get("status") not in ("completed", "failed", "cancelled"):
            continue
        print("terminal:", job_id, job["status"])
        if push:
            print("redeliver:", call("POST", f"/v1/jobs/{job_id}/webhook/redeliver")[0])
        open_jobs.pop(job_id)

reconcile({"job_replace_me": time.time() - 7200})

Choosing a grace period

The grace period should be longer than the slowest job you expect plus the retry window. Sume retries at a fixed 30-second spacing by default across ten attempts, so the retry window is about four and a half minutes of waiting plus the attempt time. A job that takes 8 minutes plus that window is covered by a grace of 15 minutes. The default of 600 seconds in the code is an example, so measure your own slowest job and add margin.

Run the sweeper on a schedule, for example every five minutes. Because reads have their own, larger budget, it will not touch your submit budget, even when the open list is long.

Notes

  • Redeliver cannot change the destination. A new URL means a new job.
  • Use job_id as the idempotency key in the handler, since a redeliver is a repeat.
  • Check the webhook_delivery fields on a job to see attempts and the last error. A turn that reads another member's job sees them as null.
  • If your receiver was down, fix it first. Otherwise the redelivery fails too.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume