A/B a Format run on Haiku 5.5 and Astra: use different keys
Same Idempotency-Key with a different body is a 409 on Sume. When the model field changes, derive the key from the body. A Python sketch that does.

Short answer
If you run the same Format twice to compare Claude Haiku 5.5 and GPT-6 Astra, give each run its own Idempotency-Key. The Format call docs say the same key with a different body returns 409 idempotency_conflict and nothing runs, while the same key and the same body returns the original receipt with idempotency_hit: true. Changing the model field changes the body.
Deriving the key from the whole body, model included, makes both cases safe: a retry reuses the key, and a variant gets a new one.
What the docs say
Statuses and codes below are from the Format call and errors pages.
| You send | Result |
|---|---|
| New key | 202 and a fresh run |
| Same key, same body | 200, original run, idempotency_hit: true |
| Same key, different body (different model, instruction or attachments) | 409 idempotency_conflict, nothing runs |
| Concurrent duplicate | 409 idempotency_key_in_use may appear |
Sketch
The function below hashes the body and posts it. It refuses to run without an API key. Model ids come from the registry spellings; send only ids your workspace accepts.
import hashlib, json, os, urllib.request
def run(handle, slug, model, cap=5):
key = os.environ.get("SUME_API_KEY", "")
if not key:
raise SystemExit("SUME_API_KEY is empty")
body = {"model": model, "generation_spend_cap_usd": cap}
raw = json.dumps(body, sort_keys=True).encode()
idem = hashlib.sha256(raw).hexdigest()[:32]
req = urllib.request.Request(
f"https://api.sume.com/v1/formats/{handle}/{slug}/runs",
data=raw,
headers={"Authorization": f"Bearer {key}",
"Content-Type": "application/json",
"Idempotency-Key": idem},
method="POST",
)
with urllib.request.urlopen(req) as resp:
return json.load(resp)["data"]["id"]
if __name__ == "__main__":
for m in ("claude-haiku-5.5", "gpt-6-astra"):
print(m, run("acme", "teaser", m))Compare fairly
Read usage.debited_usd_micros on both receipts, since the generation cap does not include the agent's own tokens. Anthropic lists Haiku 5.5 default effort as medium and OpenAI states no default for Astra, so the two are not run at the same depth by default. Note that and keep the same instruction and inputs.
Reading the two results
Put the two receipts in one table: model, debited_usd_micros, status, duration_ms if present, and a human score for the output. A single pair proves little, so run several pairs with different inputs. Keep each pair's keys different by design, and keep each retry's key the same by design; the hash does both.
If one side hits the spend cap and the other does not, the cap is doing its job and tells you which model plans longer sequences. The cap counts media spend only, so a hit means more media was requested, not more tokens.
A fair test also fixes everything except the model. Use the same Format, the same input, the same cap and the same attachments on both sides, and send both requests close together so any change to the Format itself does not land between them. Record the Format version if the response shows one. Without those controls, a difference in output may come from the input and not the model, and a difference in cost may come from a longer prompt and not a pricier rate card. Keep the sample honest: say how many pairs you ran, and do not generalize from two or three.
Finally, decide the winning rule before you look. For example: pick the cheaper model unless a reviewer scores the other clearly higher on at least most pairs. Writing the rule first stops the result from steering the rule.
Keep the results table small and dated. A row per pair, with the model, the debited amount and your score, is enough for a team to see whether a pattern holds.
- Same key, same body: replay.
- Different model: new key, new run.
Sources
Related posts
More in Developers
- AbortSignal.any: stop a Sume request on SIGTERM or after 5 seconds
Combine a shutdown signal and AbortSignal.timeout with AbortSignal.any, then tell a timeout from a SIGTERM stop on a Sume GET /v1/me call in Node.
- Retry a timed-out Sume Agent Completion without a second run
Send an Idempotency-Key on POST /v1/agent/completions. A retry returns the original receipt with idempotency_hit true; a changed payload returns a 409.
- Agent Completions 403 insufficient_scope: old key or service account?
A 403 insufficient_scope on POST /v1/agent/completions has two causes: a key made before the feature, or a service-account key. details.reason tells which.
- Agent Completions model field is sume-agent only: no Haiku or GLM
You cannot choose Claude Haiku 5.5 or GLM 5.3 in the Agent Completions model field. Sume accepts sume-agent and returns 400 for anything else.
Written by Sume