Resume a bulk run after a crash: replay the Idempotency-Key

If your client dies after POSTing a 100-item bulk run, replay the same key and same body: Sume returns 202 with the original queue. A new body gets 409.

4 min readSume
All posts

Replay the exact same request with the same Idempotency-Key. Sume answers 202 with the queue that already exists, so a crashed client can recover without creating a second batch or paying twice. Change the payload under the same key and you get 409 idempotency_conflict, with details.queue_id naming the original.

The Sume rules

A bulk run is POST /v1/formats/{handle}/{slug}/bulk-runs with concurrency from 1 to 16 and 1 to 100 items. The key may be a header or the body field idempotency_key; the header wins, and the scope of a key is one Format. Unlike a single run, a bulk replay stays 202, and the queue object has no idempotency_hit field.

Mint a fresh key for each new batch. A spent key returns the old queue, which is the opposite of what you want for a new list.

Bulk run idempotency, from the Sume docs read 2026-10-08
ReplayResult
Same key, same { concurrency, items }202 and the existing queue
Same key, different payload409 idempotency_conflict
No key, retriedA second queue
Items per queue1 to 100
Concurrency window1 to 16

How fal does it

The fal queue page, read today, says to use the request_id for idempotency and that requests in the queue are never dropped. It also says fal retries on 503 and 504 up to 10 times unless you send X-Fal-No-Retry. The idea is the same: one stable id for one unit of work. On Sume the id is a key you choose before you send, so you can store it first.

Persist the key before you send

Write the key and the item list to disk, then post. After a crash, read both back and send them unchanged.

import json, os, urllib.request, uuid

def create(items: list, state="batch.json") -> dict:
    if os.path.exists(state):
        saved = json.load(open(state))
    else:
        saved = {"key": str(uuid.uuid4()), "items": items}
        json.dump(saved, open(state, "w"))
    req = urllib.request.Request(
        "https://api.sume.com/v1/formats/acme/product-promo/bulk-runs",
        data=json.dumps({"concurrency": 4, "items": saved["items"]}).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json",
                 "Idempotency-Key": saved["key"]},
    )
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)["data"]

After the replay

Poll GET /v1/format-run-queues/{id} (formats:read). A queue status of completed means every item is terminal, not that all succeeded; branch on counts.failed. There is no queue-level webhook, only a per-item communication.webhook_url.

What the queue does while you were away

The queue keeps running with no client attached. Its controller holds a window of concurrency children in flight, and starts the next one as each ends. So after a crash that lasts minutes, you will often find some items finished. Poll the queue once to see counts, and use each item status and run_id to reconcile with your own store.

Each child can have its own communication.webhook_url, so you may already have received events for the finished ones. Dedupe on request_id for run webhooks, since the body has no job_id.

Do not reuse a key for a changed list. If you add a row to the spreadsheet and replay, you get a 409 and nothing is added. Create a new queue with a new key for the new rows only. Also, do not generate the key inside the retry loop; a key made fresh on each attempt is no key at all, and each attempt creates a new queue that is billed.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume