Two workers, one order: Idempotency-Key from order id and version
Two queue workers pick up the same order and both submit to Sume. Build the key from order id plus version so duplicates collapse and edits still create a run.

Derive the Idempotency-Key from your order id and a version number you increment whenever the order's content changes, for example order-77-v3. Two workers that pick up the same order then send the same key and the same body, and Sume returns the one run. A customer who edits the order gets version 4, a new key, and a new run. A random UUID per attempt does the opposite: it makes every worker retry, every redelivery and every crashed-and-restarted process a fresh paid run.
What Sume does with the key
The Format call docs describe a header of up to 255 characters, scoped per Format. Same key and same body: 200 with idempotency_hit: true and the existing run. Same key, different body: 409 idempotency_conflict. Same key while the first create is still in flight: 409 idempotency_key_in_use, retryable after about a second. If the first create failed with 402 or 503, the key is released, so the retry is a real attempt.
Two details shape the design. A bulk create that replays keeps returning 202 with the old queue. And a failed run does not get retried by re-sending its key, since that replays the failed run; retry with a new key or continue with previous_run_id.
Simulating the race
The script below fakes Sume's key store in memory to show the four cases: first submit, a duplicate from a second worker, an edited body that forgot to bump the version, and the edit shipped properly. It needs only Python 3. The conflict line is the useful one: it is your safety net when a developer changes the body without touching the version, and you want that to be loud rather than silent.
import json
store = {} # stands in for Sume's per-Format key store
def create(key, body):
if key in store:
saved_key_body, run = store[key]
if saved_key_body != body:
return 409, "idempotency_conflict"
return 200, run + " (idempotency_hit)"
store[key] = (body, f"arun_{len(store) + 1}")
return 202, store[key][1]
def key_for(order_id, version):
return f"order-{order_id}-v{version}"
body_v1 = json.dumps({"instruction": "15s teaser", "n": 1}, sort_keys=True)
body_v2 = json.dumps({"instruction": "15s teaser, new price", "n": 1}, sort_keys=True)
print(create(key_for(77, 1), body_v1)) # worker A
print(create(key_for(77, 1), body_v1)) # worker B, same order
print(create(key_for(77, 1), body_v2)) # edited body, forgot to bump
print(create(key_for(77, 2), body_v2)) # edit shipped as version 2What belongs in the version
Version should move whenever the paid result should differ. That includes the instruction, the input object, the attachments list and the model. It should not move for retries, redeliveries or queue replays, because those are exactly the duplicates you want collapsed. Store the version next to the order and the returned run id in the same transaction as the submit if you can, so a restart reads the row instead of guessing.
| Event | Version | Key | Result |
|---|---|---|---|
| Worker crash and restart | Same version | Same key | Replay, one run |
| Queue redelivers the job | Same version | Same key | Replay, one run |
| Customer edits the brief | Version + 1 | New key | New run, new bill |
| Retry after a failed run | Same version | New key (or previous_run_id) | New attempt, deliberate |
| Body changed, version not | Same version | Same key | 409 idempotency_conflict |
Two cautions
Do not put secrets or customer emails in the key; it is an identifier that can appear in logs. And remember that the key is scoped per Format, so the same string on two Formats is two independent keys. If you need to find runs created by a crashed process, the order-and-version key makes them easy to reconstruct, as in the crash recovery post.
Storing the result
Write the run id back against the order and version in the same transaction that records the submit, and make that write idempotent too. If two workers race, one wins the insert and the other reads it. When the replay response says idempotency_hit: true, treat it exactly like a fresh 202: store the run id, start polling or wait for the webhook. The only difference is that you did not pay twice, and your logs should say so, so you can see how often your fleet duplicates work.
Sources
Related posts
More in Developers
- Undo for AI image edits: keep a version chain of every saved result
Generative edits have no undo button. Download each result, hash it, record its parent and prompt, and walk the chain back. Python, no database.
- unsupported_media_type: video_url served as text/html or an image
Sume video trim and filter HEAD the source and refuse a declared non-video content type. What is checked, why octet-stream passes, and a runnable check.
- Valibot safeParse on a Sume job status: keep polling on bad data
Valibot's safeParse returns a result instead of throwing, so a malformed Sume status body can be logged and retried. A short schema and runnable code.
- Validate video duration and resolution in Python before you submit
Fetch GET /v1/videos/models and check duration, resolution and aspect_ratio per model in about 25 lines of Python, before a Sume video job fails.
Written by Sume