Idempotency-Key for a SaaS: customer, order and version
Derive a Format run's Idempotency-Key from customer id, order id, Format slug and a version you bump on purpose, so double clicks never make a second paid run.

Hash your tenant id, your order id, the Format slug and a version string you control, and send the result as Idempotency-Key. The key then identifies the thing being made rather than the moment you asked, so a double click or a redelivered job returns the original run instead of starting a second paid one.
A random UUID per request defeats this. It makes the header decorative, because every retry looks like a new run.
A key derived from what is being made
Namespace by customer. A key built only from the order id lets two tenants with colliding ids share a run. Include a version you bump when you deliberately want a re-run of the same order, for example after the customer edits the brief. Keep it short: the sample below truncates a SHA-256 to 40 hex characters.
import hashlib
def run_key(customer_id: str, order_id: str, version: int = 1) -> str:
raw = f"{customer_id}:{order_id}:product-promo:v{version}"
return hashlib.sha256(raw.encode()).hexdigest()[:40]
print(run_key("cust_42", "order_9001"))
print(run_key("cust_42", "order_9001", version=2))
What a replay returns
The behavior on the wire is exact, and your handler should expect all three outcomes. A conflict means your key derivation is unstable, so fix that before retrying rather than adding a retry loop.
| Replay | Result |
|---|---|
| Same key, same body | 200 with the original receipt and idempotency_hit: true, no second run and no second charge |
| Same key, different body, including a different instruction | 409 idempotency_conflict, nothing runs |
| No key | Every call starts a new paid run |
Store the run id before you answer the browser
Write the returned run_id against your record before you respond. You can re-derive the key later to find the run, but a stored id is one lookup instead of one replay. If a run fails, retry with a new key: the old one is bound to the receipt you already have, and reusing it returns that same failed receipt.
For a batch, mint a fresh key per batch. A bulk replay of a spent key returns 202 with the old queue, and keys are scoped to one Format.
Use the spend cap to express plan tiers
The same request is the place to set generation_spend_cap_usd. The cookbook sketches a mapping where a free plan gets a small cap, a pro plan a larger one, and enterprise inherits the Format's own cap. Those numbers are an example from the docs, not a recommendation; pick yours from what a run of your Format really spends.
Remember that result URLs are durable and public to anyone holding them. If one customer must never see another's output, copy the media on the webhook or proxy it through your own authenticated route before you mark the order ready.
Sources
Related posts
More in Developers
- Ideogram 4 download: Hugging Face gate, login and first image
To run Ideogram 4 locally: accept the gate on Hugging Face, log in with hf, pip install the repo, run run_inference.py. The flags and the nf4 or fp8 choice.
- Image batch in Python: read ratelimit headers and retry-after
Sume can send ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after on image calls. Back off on 429 in Python, and treat queue_full separately.
- Image edit returns 415 on Sume: send JSON, not multipart
A 415 unsupported_media_type from Sume's image API means the body was not application/json. Send reference images as public HTTPS URLs inside a JSON body.
- Image model missing from Sume's /v1/images/models? Read it live
GET /v1/images/models lists models Sume can serve now; sume/auto is never listed, and unknown ids return 404 model_not_found. Read the catalog at runtime.
Written by Sume