Reconcile Sume jobs after a deploy or outage: poll what is open
After downtime, read status for every job your own table still shows as open, honor terminal and result_ready, and never resubmit. Python with sqlite.

Do not resubmit anything. After a deploy or an outage, select the jobs your own database still shows as open, read GET /v1/jobs/{id}/status for each one, and write the terminal state when terminal is true. Jobs keep running and finishing while your receiver is down, and webhooks that failed are only a missed notification, so the status endpoint is the source of truth you recover from.
What survives an outage
Almost everything survives, because the job lives on Sume's side and not in your process. What you lose is the notification. The table separates what Sume keeps from what you have to rebuild.
| Thing | Survives? | Recovery |
|---|---|---|
| The job and its result | Yes | GET /v1/jobs/{id}/status, then /result when result_ready |
| Webhook attempts | Up to 10, at a 30 s spacing for jobs, then exhausted | Poll, or POST /v1/jobs/{id}/webhook/redeliver |
| Your in-flight HTTP waits | No | Read status; a client timeout never cancels |
| Your idempotency keys | Only if you stored them | Resend with the stored key and get the original job |
| The job list | Visible to the key's own member | GET /v1/jobs, but only for jobs that key's member created |
The reconciler
The script below reads every row whose state is submitted or working, fetches its stored status URL, and moves a row forward only when the payload is terminal. A network error, a 429 or a 5xx leaves the row alone for the next run. The final update names the allowed previous states, so a webhook that arrives at the same moment and has already written done is never overwritten. The fake responses at the bottom make it runnable without a key; swap them for the real call in production.
import json, sqlite3, urllib.request
def fetch_status(url, key):
req = urllib.request.Request(url, headers={"x-api-key": key})
with urllib.request.urlopen(req, timeout=10) as r:
return json.load(r)
def reconcile(db, key, fetch=fetch_status):
rows = db.execute("SELECT id, status_url FROM jobs WHERE state IN ('submitted','working')").fetchall()
for job_id, url in rows:
try:
s = fetch(url, key)
except Exception as exc: # network, 429, 5xx: leave the row, try next run
print(job_id, "skip", exc)
continue
if s.get("terminal"):
state = "done" if s.get("result_ready") else "failed"
db.execute("UPDATE jobs SET state=? WHERE id=? AND state IN ('submitted','working')", (state, job_id))
else:
db.execute("UPDATE jobs SET state='working' WHERE id=? AND state='submitted'", (job_id,))
db.commit()
db = sqlite3.connect(":memory:")
db.execute("CREATE TABLE jobs (id TEXT PRIMARY KEY, status_url TEXT, state TEXT)")
db.executemany("INSERT INTO jobs VALUES (?,?,?)", [("a", "u/a", "submitted"), ("b", "u/b", "working"), ("c", "u/c", "working")])
fake = {"u/a": {"terminal": True, "result_ready": True}, "u/b": {"terminal": False}, "u/c": {"terminal": True, "result_ready": False}}
reconcile(db, "test", lambda url, key: fake[url])
print(db.execute("SELECT id, state FROM jobs ORDER BY id").fetchall())Keep it cheap
Reads have their own budget, forty times the write budget for your plan, so a reconciler over a few hundred jobs fits easily. Still, do not run it as a tight loop. When a non-terminal status carries next_poll_after_seconds, store it with the row and skip that row until the time has passed. A recovery job that is too eager can turn into a read 429, which delays the recovery it was meant to perform. A short fixed delay between requests, or a small worker pool, is plenty for a job that runs once.
Run it once at startup, so that a deploy heals itself, and then on a schedule that is longer than your usual job duration. If you receive webhooks, the scheduled run is a safety net and almost always finds nothing.
Log what the reconciler changes, not only what it reads. A line such as the job id, the state it moved from and the state it moved to gives you a record of every notification that the webhook path missed, and a count of those lines per week tells you whether your receiver is healthy or quietly dropping events. Alert on a sudden jump, since it usually means a deploy broke the receiver and the reconciler is quietly covering for it, and fix the receiver before the safety net is the only thing working.
Mind the rows with no job yet
The hardest case is a crash between the submit and the database write. The job exists, but your table has no row for it. Because your retry reuses the same Idempotency-Key, resubmitting returns the original job and writes the row, with no second charge. That is why the key should be derived from your own order or intent, not generated in memory.
Rows that stay open far longer than expected deserve a human look. For a Format run, the run's expires_at is 90 minutes after creation, so a row older than that for a run is not waiting on anything. For a plain job, read the status, then the events, and decide whether to cancel, which only works before generation starts.
Sources
Related posts
More in Developers
- Redact faces and license plates: Pillow first, AI edit only to replace
For redaction use Pillow boxes you control; use an AI mask edit on openai/gpt-image-2.5 only to replace a plate or face, from $0.0094 per image on Sume.
- Rotate the Sume webhook signing secret without dropping a delivery
Upgrade the verifier first, rotate with POST /v1/webhooks/signing-secret/rotate, deploy the new secret inside the 24-hour two-signature window, then confirm it.
- Should my backend call Sume over hosted MCP or the REST API?
REST from a backend, hosted MCP from an agent client. Where they differ: auth, wait limits, REST-only Image 1.0 and Video 1.0, and write budgets.
- Speech to text API in Go: transcribe audio with net/http
Transcribe audio in Go using only the standard library: submit to Sume STT, poll the job and print the text. A 30-line program at one cent per audio minute.
Written by Sume