Lost video callback? Sweep pending jobs after the 10-attempt window
Sume retries a job webhook 10 times, 30 seconds apart. If none gets through, the job is still done. A 30-line sweeper that polls pending jobs recovers it.

If every delivery of a callback_url webhook fails, Sume stops after 10 attempts spaced 30 seconds apart, which is 270 seconds of gaps between the first and last attempt, and the job is still in its real terminal state. To recover it, poll the polling_url that you saved at submit time. A short sweeper that polls only the jobs with no callback yet is enough.
This matters for a re-render of a stored prompt library because a deploy, a certificate problem or a firewall change on the receiver loses callbacks silently, and the clips you paid for are sitting finished.
What the docs promise
The webhook docs state that Sume retries network errors and non-2xx responses up to 10 attempts, with a fixed 30-second delay by default and a 10-second timeout per attempt. After ten refused attempts you have a failed delivery and a job that still completed. The docs say delivery is an optimization and that status_url polls should stay available.
They also describe a manual path: POST /v1/jobs/{job_id}/webhook/redeliver re-sends the real terminal event with a fresh timestamp and signature, and does not count against the automatic ten.
| Option | How | Good for |
|---|---|---|
| Poll the saved polling_url | GET it with your key; read status and unsigned_urls | A few jobs, or any job with no callback yet |
| Redeliver the webhook | POST /v1/jobs/{job_id}/webhook/redeliver (jobs:write) | Replaying into your existing receiver |
| Sweep a pending table | Loop over open jobs on a timer | A whole library with a long-running worker |
A sweeper
Keep a table or JSON file of jobs that have no callback recorded, as {job_id: polling_url}. The function below polls each URL once and returns the jobs that reached completed, failed or cancelled, removing them from the pending set. Run it from cron every few minutes. Set SUME_API_KEY first.
import json
import os
import urllib.request
KEY = os.environ["SUME_API_KEY"]
DONE = {"completed", "failed", "cancelled"}
def poll(url):
req = urllib.request.Request(url, headers={"Authorization": f"Bearer {KEY}"})
with urllib.request.urlopen(req, timeout=30) as resp:
return json.load(resp)
def sweep(pending):
"""pending: {job_id: polling_url} for jobs with no callback yet."""
finished = {}
for job_id, url in list(pending.items()):
job = poll(url)
if job["status"] in DONE:
finished[job_id] = job
del pending[job_id]
return finished
if __name__ == "__main__":
waiting = json.load(open("pending.json"))
for job_id, job in sweep(waiting).items():
print(job_id, job["status"], job.get("unsigned_urls", [])[:1])
json.dump(waiting, open("pending.json", "w"))
Why a lost callback is not a lost job
The terminal state belongs to the job, not to the delivery. A job that reached completed stays completed whether or not your receiver heard about it, and unsigned_urls stays on the poll response. The cost is captured when the job completes, so a lost callback does not change the bill; it only delays when your system learns about the clip.
That is why the sweeper keys on job_id. A manual redeliver or a late retry can arrive after the sweeper has already processed the job, and the receiver should accept it and do nothing.
Timing the sweep
Start sweeping a job only after the callback window has passed, which is about 5 minutes after the job should have finished. A sweep that begins too early duplicates the callback you are about to receive and wastes requests.
When the sweeper finds a completed job, process it with the same code path the callback uses, and key your writes on job_id so a late callback that arrives afterwards does nothing. Download unsigned_urls[0] promptly and store the file in your own storage.
- Record the polling_url at submit time; it is in the 202 response.
- Treat
failedas final for the job, and resubmit with a new idempotency key if you want another take. - Keep the sweep interval above the 30-second poll the docs recommend.
Sources
Related posts
More in Developers
- callback_url on /v1/videos: the Sume job envelope that arrives
A /v1/videos callback_url delivers Sume's job.completed, job.failed or job.canceled envelope, not video.generation.* events. Payload, signature, checks.
- callback_url or webhook_url? Which Sume route takes which field name
POST /v1/videos takes callback_url; the model endpoints take mode webhook with webhook_url. Both must be public HTTPS and both deliver signed job events.
- Cancel every queued Sume job without cancelling a running one
A Python script that lists queued jobs, checks the cancelable flag, posts the cancel, and handles 409 job_generation_already_started when the race is lost.
- Cancel queued Sume jobs after queue_full; handle 409 already started
How to free capacity after a 429 queue_full: cancel queued jobs, read job_generation_already_started on running ones, and why a cancel releases the reserve.
Written by Sume