Resume a bulk run after a crash: replay the Idempotency-Key
If your client dies after POSTing a 100-item bulk run, replay the same key and same body: Sume returns 202 with the original queue. A new body gets 409.

Replay the exact same request with the same Idempotency-Key. Sume answers 202 with the queue that already exists, so a crashed client can recover without creating a second batch or paying twice. Change the payload under the same key and you get 409 idempotency_conflict, with details.queue_id naming the original.
The Sume rules
A bulk run is POST /v1/formats/{handle}/{slug}/bulk-runs with concurrency from 1 to 16 and 1 to 100 items. The key may be a header or the body field idempotency_key; the header wins, and the scope of a key is one Format. Unlike a single run, a bulk replay stays 202, and the queue object has no idempotency_hit field.
Mint a fresh key for each new batch. A spent key returns the old queue, which is the opposite of what you want for a new list.
| Replay | Result |
|---|---|
Same key, same { concurrency, items } | 202 and the existing queue |
| Same key, different payload | 409 idempotency_conflict |
| No key, retried | A second queue |
| Items per queue | 1 to 100 |
| Concurrency window | 1 to 16 |
How fal does it
The fal queue page, read today, says to use the request_id for idempotency and that requests in the queue are never dropped. It also says fal retries on 503 and 504 up to 10 times unless you send X-Fal-No-Retry. The idea is the same: one stable id for one unit of work. On Sume the id is a key you choose before you send, so you can store it first.
Persist the key before you send
Write the key and the item list to disk, then post. After a crash, read both back and send them unchanged.
import json, os, urllib.request, uuid
def create(items: list, state="batch.json") -> dict:
if os.path.exists(state):
saved = json.load(open(state))
else:
saved = {"key": str(uuid.uuid4()), "items": items}
json.dump(saved, open(state, "w"))
req = urllib.request.Request(
"https://api.sume.com/v1/formats/acme/product-promo/bulk-runs",
data=json.dumps({"concurrency": 4, "items": saved["items"]}).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": saved["key"]},
)
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)["data"]After the replay
Poll GET /v1/format-run-queues/{id} (formats:read). A queue status of completed means every item is terminal, not that all succeeded; branch on counts.failed. There is no queue-level webhook, only a per-item communication.webhook_url.
What the queue does while you were away
The queue keeps running with no client attached. Its controller holds a window of concurrency children in flight, and starts the next one as each ends. So after a crash that lasts minutes, you will often find some items finished. Poll the queue once to see counts, and use each item status and run_id to reconcile with your own store.
Each child can have its own communication.webhook_url, so you may already have received events for the finished ones. Dedupe on request_id for run webhooks, since the body has no job_id.
Do not reuse a key for a changed list. If you add a row to the spreadsheet and replay, you get a 409 and nothing is added. Create a new queue with a new key for the new rows only. Also, do not generate the key inside the retry loop; a key made fresh on each attempt is no key at all, and each attempt creates a new queue that is billed.
Sources
Related posts
More in Developers
- Sume Image API 400 unsupported_parameter: which field fails where
Which Sume image models list quality, resolution, mask_url, background, output_format and references, and which never do (seed, stream). Plus a check script.
- result_ready vs terminal vs completed: gate the Sume result fetch
Poll a Sume job until terminal, fetch the result only when result_ready is true. Failed and canceled jobs answer 409 job_not_completed on /result.
- Sume TTS word timestamps to caption cues for a narrated 60-second clip
Ask Sume TTS for timestamps.words, group them into cues and send them to video-captions so no recognition runs. About 25 cents for a 60-second narration.
- Sume TypeScript SDK waitForJob: 20-minute timeout, job keeps billing
How @sume-com/sdk waitForJob polls a generation job, what SumeJobTimeoutError means, and why a client timeout does not cancel or refund the job.
Written by Sume