Crash-safe Sume submit: write the key first, reconcile on boot

If a worker dies between POST and saving the job id, list queued and processing jobs, match idempotency_key, and resubmit only keys still unknown.

4 min readSume
All posts

Write your Idempotency-Key to your own database before you call Sume, then on startup list the workspace's queued and processing jobs and match each row's idempotency_key back to your keys. Any key with no matching job is safe to resubmit with the same key.

This closes the gap that a plain retry loop cannot: the crash after Sume accepted the job but before your code saved its id. The GET /v1/jobs response includes idempotency_key on every row for exactly this join, per the OpenAPI description read 2026-10-10.

The order of operations

  • Insert a row (key, job_id NULL) in your own store and commit it.
  • POST the job with that key in the Idempotency-Key header.
  • On success, update the row with the returned job id.
  • On boot, run the reconcile pass below, then resubmit the keys that are still null.

A reconcile pass

The loop follows data.next_cursor back as starting_after. The cursor is absent on the last page, which is the loop's end; do not build one from the final job. It uses urllib with a custom User-Agent, since the default agent was refused in my 403 post test.

import json, os, sqlite3, urllib.parse, urllib.request
H = {"x-api-key": os.environ["SUME_API_KEY"], "User-Agent": "recover/1.0"}
db = sqlite3.connect("orders.db")
db.execute("create table if not exists sub (key text primary key, job_id text)")

def active_jobs(status):
    cursor = None
    while True:
        q = {"status": status, "limit": 100}
        if cursor:
            q["starting_after"] = cursor
        url = "https://api.sume.com/v1/jobs?" + urllib.parse.urlencode(q)
        with urllib.request.urlopen(urllib.request.Request(url, headers=H), timeout=30) as r:
            page = json.load(r)["data"]
        yield from page["jobs"]
        cursor = page.get("next_cursor")
        if not cursor:
            return

def reconcile():
    for status in ("queued", "processing"):
        for job in active_jobs(status):
            db.execute("update sub set job_id = ? where key = ? and job_id is null",
                       (job["id"], job["idempotency_key"]))
    db.commit()
    return [k for (k,) in db.execute("select key from sub where job_id is null")]

print("still unknown:", reconcile())

Limits of this approach

The pass only lists queued and processing. A job that already finished during the outage will not show up there, so for those keys the resubmit relies on the same-key rule in the Sume docs: use the same key again only for the same operation and payload. If you want certainty before spending, also scan completed and failed for recent pages, or look up your keys in the usage ledger.

List pages are newest first, which is not the order you submitted in. Always match on idempotency_key, never on position. Keep the sqlite file on durable storage, or the whole scheme starts from nothing.

Every crash point, and what recovers it

A crash can land in four places, and only one of them needs the list call. The table covers each one so you can test them by killing the process on purpose.

Crash points in the submit flow, reasoning from the Sume idempotency docs read 2026-10-10
Crash happensState on diskWhat recovers
Before your insertNo rowNothing was started; the item is still in your queue
After insert, before POSTKey, job id nullResubmit with the stored key
After POST, before saving the idKey, job id nullReconcile finds the job by idempotency_key
After saving the idKey and job idResume polling the job id

Sources

Related posts

More in Developers

All Developers posts

Written by Sume