Agent crashed after submitting: find the Sume runs it started

List a Format's runs, read trigger.idempotency_key, and rebuild which rows already ran after a fresh agent session or crash, without paying for a duplicate.

5 min readSume
All posts

Call GET /v1/formats/{handle}/{slug}/runs, page through it, and read trigger.idempotency_key on each run. If you named every submit with a key you can rebuild, such as one per row, that list tells you which rows already started and what state each is in, even though the process that submitted them has lost its memory.

This matters more now that agents are built to lose context on purpose. Hermes starts every cron job in a fresh session, and Cloudflare's new Pi harness for its Agents SDK is described as persisting work through interruption, which means an agent can resume and run a step again.

Why would a resumed agent submit twice?

Cloudflare's changelog says PiHarness is in beta and Pi Durable is experimental, and its example declares a tool with replay: "safe". I did not verify, and do not claim, how recovery decides which tools to run again. The point for a paid API is general: any tool that creates a Format run can be called a second time after a resume, so its second call must be harmless. A stable Idempotency-Key makes it harmless. Listing runs is the second line of defense, for when the key was never recorded or the payload changed.

What does the list give you?

There is no list across Formats, so you search the one Format you submitted to, and there is no GET /v1/format-runs either.

List runs for a Format, read 2026-10-03
PropertyValue
PathGET /v1/formats/{handle}/{slug}/runs
OrderNewest first
Page sizelimit from 1 to 100, default 20
PagingPass next_cursor back as cursor until has_more is false
Per runid, status, created_at and trigger.idempotency_key
Scopeformats:read
Never runReturns an empty list, not a 404

How do I rebuild the state?

This reads a few pages and indexes runs by key. Cap the pages, because you only need the window since the job started.

import os, requests

URL = "https://api.sume.com/v1/formats/myteam/product-promo/runs"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def runs_by_key(max_pages=5):
    out, cursor = {}, None
    for _ in range(max_pages):
        params = {"limit": 100}
        if cursor:
            params["cursor"] = cursor
        r = requests.get(URL, params=params, headers=H, timeout=30)
        r.raise_for_status()
        page = r.json()
        for run in page["data"]:
            key = (run.get("trigger") or {}).get("idempotency_key")
            if key:
                out[key] = run
        if not page["has_more"]:
            break
        cursor = page["next_cursor"]
    return out

seen = runs_by_key()
for row_id in ["row-1", "row-2", "row-3"]:
    run = seen.get(row_id)
    print(row_id, run["status"] if run else "not started")

What do you do with the answer?

Reconcile once, then submit only what is missing.

  • Rows that show completed need no new run. Read their output from the receipt and move on.
  • Rows that show queued or processing are still working; wait for the webhook or poll status_url.
  • Rows that show failed need a new key, because replaying the old one returns the failed run. Read the error first; an unattended_blocked failure needs a changed input.
  • Rows that are not started can be submitted now with the key you planned.

What does this not cover?

A bulk queue's children are a different matter. I could not confirm from the docs which idempotency key a bulk child records, and the queue id is not listable, so if you may need to find rows after a crash, store the frq_ id at the moment of the 202 or submit single runs with row keys. Also keep listings short: the cursor is stable while runs are created, but a long history means more pages, and every page is a read against a budget that is forty times the write budget.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume