Recover Sume jobs after a crash: match your key to GET /v1/jobs

Your worker died after submit and lost the job ids. Page GET /v1/jobs, match idempotency_key to your own keys, stop at the last page. A 29-line sample.

4 min readSume
All posts

Page through GET /v1/jobs and match each job's idempotency_key against the keys you stored before submitting; the match gives you the job id you lost. The OpenAPI description of that field calls it the join column for recovery, and it explains why: list pages are newest-first, which is not your submit order once a wave retries, so you map your key to id and never rely on array position.

If you still hold the original request body, there is an even simpler fix: submit again with the same key, and Sume returns the original job instead of a second one. The list route is for when you want to reconcile many intents without re-sending bodies.

The loop

The list response is data.jobs plus an optional data.next_cursor. The schema says the cursor is opaque, present only when more jobs remain, and absent on the last page, which is the loop terminator: do not synthesize one from the final job. The sample stops as soon as every wanted key is found, or when the cursor is gone. I ran it against a local stand-in that served two pages.

import json, os, sys, urllib.parse, urllib.request

BASE = "https://api.sume.com"


def get(path: str) -> dict:
    req = urllib.request.Request(BASE + path, headers={"x-api-key": os.environ["SUME_API_KEY"]})
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)["data"]


def jobs_by_key(wanted: set[str]) -> dict[str, str]:
    """Map your Idempotency-Keys to job ids; pages are newest first."""
    found, cursor = {}, None
    while True:
        q = {"limit": 100, **({"starting_after": cursor} if cursor else {})}
        page = get("/v1/jobs?" + urllib.parse.urlencode(q))
        for job in page["jobs"]:
            if job.get("idempotency_key") in wanted:
                found[job["idempotency_key"]] = job["id"]
        cursor = page.get("next_cursor")  # absent on the last page
        if not cursor or wanted <= found.keys():
            return found


keys = set(sys.argv[1:])
found = jobs_by_key(keys)
for k in sorted(keys):
    print(k, "->", found.get(k, "not listed"))

What each parameter does (OpenAPI read 2026-10-10)

GET /v1/jobs parameters from the reference JSON, read 2026-10-10
ParameterValuesUse in recovery
limit1 to 100Use 100 to cut the page count
statusqueued, processing, completed, failed, canceledFilter to queued or processing to find work in flight
typeany non-empty stringNarrow to one route family if you know it
starting_afteropaque cursorPass back next_cursor unchanged

Rules for a safe reconcile

  • Store the Idempotency-Key in your own database before the first send, so there is something to match after a crash.
  • Treat a key with no match as unknown, not as proof the job does not exist. Resubmit with the same key, and Sume returns the original job if there was one.
  • Never build a new key for the resubmit. A new key is a new paid job.
  • Read the job's status and terminal after matching, then continue the normal poll or webhook flow.
  • The list is a read, so it spends the read bucket, not the write bucket; page size 100 keeps it cheap.

When the list is too slow

A workspace with many jobs makes a long walk, and a walk that goes back days to find a one-hour-old key is wasted work. Because pages are newest first, a recent crash is found in the first page or two. If you record the submit time with each key, you can also stop once a page's oldest job is older than your oldest unmatched submit.

For a very large fan-out, record keys in a table with a state column and run the reconcile as a scheduled job rather than at startup, so a restart storm does not become a read storm.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume