Lost video callback? Sweep pending jobs after the 10-attempt window

Sume retries a job webhook 10 times, 30 seconds apart. If none gets through, the job is still done. A 30-line sweeper that polls pending jobs recovers it.

5 min readSume
All posts

If every delivery of a callback_url webhook fails, Sume stops after 10 attempts spaced 30 seconds apart, which is 270 seconds of gaps between the first and last attempt, and the job is still in its real terminal state. To recover it, poll the polling_url that you saved at submit time. A short sweeper that polls only the jobs with no callback yet is enough.

This matters for a re-render of a stored prompt library because a deploy, a certificate problem or a firewall change on the receiver loses callbacks silently, and the clips you paid for are sitting finished.

What the docs promise

The webhook docs state that Sume retries network errors and non-2xx responses up to 10 attempts, with a fixed 30-second delay by default and a 10-second timeout per attempt. After ten refused attempts you have a failed delivery and a job that still completed. The docs say delivery is an optimization and that status_url polls should stay available.

They also describe a manual path: POST /v1/jobs/{job_id}/webhook/redeliver re-sends the real terminal event with a fresh timestamp and signature, and does not count against the automatic ten.

Recovering a lost video callback (Sume docs, as of 2026-10-08)
OptionHowGood for
Poll the saved polling_urlGET it with your key; read status and unsigned_urlsA few jobs, or any job with no callback yet
Redeliver the webhookPOST /v1/jobs/{job_id}/webhook/redeliver (jobs:write)Replaying into your existing receiver
Sweep a pending tableLoop over open jobs on a timerA whole library with a long-running worker

A sweeper

Keep a table or JSON file of jobs that have no callback recorded, as {job_id: polling_url}. The function below polls each URL once and returns the jobs that reached completed, failed or cancelled, removing them from the pending set. Run it from cron every few minutes. Set SUME_API_KEY first.

import json
import os
import urllib.request

KEY = os.environ["SUME_API_KEY"]
DONE = {"completed", "failed", "cancelled"}


def poll(url):
    req = urllib.request.Request(url, headers={"Authorization": f"Bearer {KEY}"})
    with urllib.request.urlopen(req, timeout=30) as resp:
        return json.load(resp)


def sweep(pending):
    """pending: {job_id: polling_url} for jobs with no callback yet."""
    finished = {}
    for job_id, url in list(pending.items()):
        job = poll(url)
        if job["status"] in DONE:
            finished[job_id] = job
            del pending[job_id]
    return finished


if __name__ == "__main__":
    waiting = json.load(open("pending.json"))
    for job_id, job in sweep(waiting).items():
        print(job_id, job["status"], job.get("unsigned_urls", [])[:1])
    json.dump(waiting, open("pending.json", "w"))

Why a lost callback is not a lost job

The terminal state belongs to the job, not to the delivery. A job that reached completed stays completed whether or not your receiver heard about it, and unsigned_urls stays on the poll response. The cost is captured when the job completes, so a lost callback does not change the bill; it only delays when your system learns about the clip.

That is why the sweeper keys on job_id. A manual redeliver or a late retry can arrive after the sweeper has already processed the job, and the receiver should accept it and do nothing.

Timing the sweep

Start sweeping a job only after the callback window has passed, which is about 5 minutes after the job should have finished. A sweep that begins too early duplicates the callback you are about to receive and wastes requests.

When the sweeper finds a completed job, process it with the same code path the callback uses, and key your writes on job_id so a late callback that arrives afterwards does nothing. Download unsigned_urls[0] promptly and store the file in your own storage.

  • Record the polling_url at submit time; it is in the 202 response.
  • Treat failed as final for the job, and resubmit with a new idempotency key if you want another take.
  • Keep the sweep interval above the 30-second poll the docs recommend.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume