Reconcile Sume jobs after a deploy or outage: poll what is open

After downtime, read status for every job your own table still shows as open, honor terminal and result_ready, and never resubmit. Python with sqlite.

5 min readSume
All posts

Do not resubmit anything. After a deploy or an outage, select the jobs your own database still shows as open, read GET /v1/jobs/{id}/status for each one, and write the terminal state when terminal is true. Jobs keep running and finishing while your receiver is down, and webhooks that failed are only a missed notification, so the status endpoint is the source of truth you recover from.

What survives an outage

Almost everything survives, because the job lives on Sume's side and not in your process. What you lose is the notification. The table separates what Sume keeps from what you have to rebuild.

What survives downtime and how to recover it (read 2026-10-07)
ThingSurvives?Recovery
The job and its resultYesGET /v1/jobs/{id}/status, then /result when result_ready
Webhook attemptsUp to 10, at a 30 s spacing for jobs, then exhaustedPoll, or POST /v1/jobs/{id}/webhook/redeliver
Your in-flight HTTP waitsNoRead status; a client timeout never cancels
Your idempotency keysOnly if you stored themResend with the stored key and get the original job
The job listVisible to the key's own memberGET /v1/jobs, but only for jobs that key's member created

The reconciler

The script below reads every row whose state is submitted or working, fetches its stored status URL, and moves a row forward only when the payload is terminal. A network error, a 429 or a 5xx leaves the row alone for the next run. The final update names the allowed previous states, so a webhook that arrives at the same moment and has already written done is never overwritten. The fake responses at the bottom make it runnable without a key; swap them for the real call in production.

import json, sqlite3, urllib.request

def fetch_status(url, key):
    req = urllib.request.Request(url, headers={"x-api-key": key})
    with urllib.request.urlopen(req, timeout=10) as r:
        return json.load(r)

def reconcile(db, key, fetch=fetch_status):
    rows = db.execute("SELECT id, status_url FROM jobs WHERE state IN ('submitted','working')").fetchall()
    for job_id, url in rows:
        try:
            s = fetch(url, key)
        except Exception as exc:  # network, 429, 5xx: leave the row, try next run
            print(job_id, "skip", exc)
            continue
        if s.get("terminal"):
            state = "done" if s.get("result_ready") else "failed"
            db.execute("UPDATE jobs SET state=? WHERE id=? AND state IN ('submitted','working')", (state, job_id))
        else:
            db.execute("UPDATE jobs SET state='working' WHERE id=? AND state='submitted'", (job_id,))
    db.commit()

db = sqlite3.connect(":memory:")
db.execute("CREATE TABLE jobs (id TEXT PRIMARY KEY, status_url TEXT, state TEXT)")
db.executemany("INSERT INTO jobs VALUES (?,?,?)", [("a", "u/a", "submitted"), ("b", "u/b", "working"), ("c", "u/c", "working")])
fake = {"u/a": {"terminal": True, "result_ready": True}, "u/b": {"terminal": False}, "u/c": {"terminal": True, "result_ready": False}}
reconcile(db, "test", lambda url, key: fake[url])
print(db.execute("SELECT id, state FROM jobs ORDER BY id").fetchall())

Keep it cheap

Reads have their own budget, forty times the write budget for your plan, so a reconciler over a few hundred jobs fits easily. Still, do not run it as a tight loop. When a non-terminal status carries next_poll_after_seconds, store it with the row and skip that row until the time has passed. A recovery job that is too eager can turn into a read 429, which delays the recovery it was meant to perform. A short fixed delay between requests, or a small worker pool, is plenty for a job that runs once.

Run it once at startup, so that a deploy heals itself, and then on a schedule that is longer than your usual job duration. If you receive webhooks, the scheduled run is a safety net and almost always finds nothing.

Log what the reconciler changes, not only what it reads. A line such as the job id, the state it moved from and the state it moved to gives you a record of every notification that the webhook path missed, and a count of those lines per week tells you whether your receiver is healthy or quietly dropping events. Alert on a sudden jump, since it usually means a deploy broke the receiver and the reconciler is quietly covering for it, and fix the receiver before the safety net is the only thing working.

Mind the rows with no job yet

The hardest case is a crash between the submit and the database write. The job exists, but your table has no row for it. Because your retry reuses the same Idempotency-Key, resubmitting returns the original job and writes the row, with no second charge. That is why the key should be derived from your own order or intent, not generated in memory.

Rows that stay open far longer than expected deserve a human look. For a Format run, the run's expires_at is 90 minutes after creation, so a row older than that for a run is not waiting on anything. For a plain job, read the status, then the events, and decide whether to cancel, which only works before generation starts.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume