SHA-256 the batch body into the Sume idempotency key for bulk chunks

Derive each bulk chunk's Idempotency-Key from a hash of its items so replays reuse the queue and edited rows get a new key instead of a 409 conflict.

5 min readSume
All posts

Hash the canonical JSON of a chunk's items and concurrency, prefix it with a campaign name, and send that as the Idempotency-Key. The same chunk then always produces the same key, so a retry returns the original queue with 202 and renders nothing twice. A chunk whose rows were edited produces a different key, which starts a new queue. The alternative, a hand-picked key such as bf26-chunk-2, fails differently: after an edit the same key meets a different body and the API answers 409 idempotency_conflict.

Neither behavior is better in the abstract. A content hash trades the loud conflict for a quiet new queue, so use it where re-rendering after an edit is what you want, and hand-picked keys where it is not.

Key and chunk functions

Rows are sorted by SKU before chunking, so the order in your spreadsheet does not move rows between chunks. The demo at the bottom builds 230 rows, shows the chunk sizes, proves a reversed input gives identical keys, and changes one title.

import hashlib, json

def batch_key(prefix, items, concurrency):
    canon = json.dumps({"c": concurrency, "i": items}, sort_keys=True,
                       separators=(",", ":"), ensure_ascii=False)
    return f"{prefix}-{hashlib.sha256(canon.encode()).hexdigest()[:32]}"

def chunks(rows, size=100):
    rows = sorted(rows, key=lambda r: r["sku"])
    return [rows[i:i + size] for i in range(0, len(rows), size)]

if __name__ == "__main__":
    rows = [{"sku": f"S{n:03}", "input": {"title": f"Item {n}"}} for n in range(230)]
    keys = [batch_key("bf26", [{"input": r["input"]} for r in c], 4) for c in chunks(rows)]
    print([len(c) for c in chunks(rows)], len(keys[0]))
    again = [batch_key("bf26", [{"input": r["input"]} for r in c], 4) for c in chunks(rows[::-1])]
    assert keys == again
    rows[0]["input"]["title"] = "Changed"
    changed = [batch_key("bf26", [{"input": r["input"]} for r in c], 4) for c in chunks(rows)]
    print([a == b for a, b in zip(keys, changed)])

What the demo shows

Run it and it prints [100, 100, 30] and a key length of 37. Then it prints [False, True, True]: editing the first row changed the key for chunk 0 only, because chunks 1 and 2 hold the same rows as before. Re-sending those two chunks would replay their existing queues.

Inserting a row instead of editing one shifts every later chunk boundary by one, and every chunk after it hashes differently. If your catalog gains rows often, key on a stable group, such as a category or a SKU range, and chunk inside it.

Key behavior of the Format API with the same key (read 2026-10-07)
RequestResult
Same key, same bodyReplay: the original queue comes back with 202, nothing new starts
Same key, different body409 idempotency_conflict, with the original details.queue_id
New key, any bodyA new queue, new runs and new charges

Rules to keep

A key is scoped to one Format and may be up to 255 characters, so a 37-character hash fits easily. The key covers the whole body, including concurrency. Changing the window from 4 to 8 on a retry is a different body and conflicts under a fixed key, but under a hash key it simply makes a new queue, which is almost never what a retry should do. Keep concurrency constant between attempts for a given chunk.

Store the key next to the queue id in your own table the moment the create call returns. If the process dies between the 202 and the write, you can rebuild the same key from the same rows and send the create again, which hands back the queue you lost instead of starting a second one. This is the only recovery path for a lost queue id, because the API has no list-queues endpoint.

Hash what you send, not what you read. Build the items first, with the same field order and the same defaults, and only then hash them. Hashing the raw spreadsheet row means a harmless change to an unused column changes the key.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume