A/B test two LLMs on one Sume Format input with the model field
Run one Sume Format twice with a different model, then compare the receipt's model and debited cost. A short Python script with an Idempotency-Key and polling.

Call the same Format twice with identical input and a different model, poll both receipts, and compare model and usage.debited_usd_micros. The script below does that in under 30 lines.
It uses the documented create a run endpoint and the receipt fields. model selects only the orchestrating LLM. The image, video and audio models still come from the Format's own tools.
What does the script look like?
The input object (product_name here) is a placeholder, so shape it to the Format's declared io.input_kind from the Format catalog. The create response and the receipt follow the shapes in the docs; if your first call returns an error, the error code in the body tells you which field to fix.
import os, time, requests
API = "https://api.sume.com/v1"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
BODY = {"instruction": "One 15s vertical ad, no captions.",
"input": {"product_name": "Aurora Headphones"}}
MODELS = ["openrouter/anthropic-claude-opus-5-5",
"openrouter/deepseek-deepseek-v4.1-flash"]
def run(model, key):
r = requests.post(f"{API}/formats/sume/sume-product-commercial/runs",
headers={**H, "Idempotency-Key": key},
json={**BODY, "model": model}, timeout=60)
r.raise_for_status()
return r.json()["data"]["id"]
def wait(run_id):
while True:
d = requests.get(f"{API}/format-runs/{run_id}", headers=H, timeout=60).json()["data"]
if d["status"] in ("completed", "failed", "canceled", "skipped"):
return d
time.sleep(15)
ids = [run(m, f"ab-test-1-{i}") for i, m in enumerate(MODELS)]
for run_id in ids:
d = wait(run_id)
print(d["model"], d["status"], (d["usage"] or {}).get("debited_usd_micros"))
What do you compare?
Debited is the number to compare, because it is the one that includes the model you changed.
model: the catalog id the orchestrator ran on. If it differs from what you sent, a retirement moved the request.usage.debited_usd_micros: what the wallet deducted, including the agent's own LLM turn. Divide by 1,000,000 for dollars.usage.billable_amount_usd_micros: generation spend only. It excludes the LLM turn, so it will look similar across models and hide the difference you are testing.status,errorand the artifacts: a cheaper run that fails is not cheaper.
How do you keep the test fair?
Use the same instruction and input in both runs, and run each pair more than once, since one sample of an agent is anecdote. Keep the Idempotency-Key distinct per run and per repeat, or the second call replays the first. Send a per-run spend cap if you want a ceiling on a test that renders video.
Do not use previous_run_id to compare. That continues one thread, and the second run would see the first's work, which is not a fair second sample.
An id outside the catalog returns 400 invalid_request, so a failed POST on one id usually means the id is wrong or no longer listed.
What does this not tell you?
It tells you cost and completion on your input. It does not grade the video. Score the outputs yourself, or with a fixed rubric, and write the score next to the receipt id. The cost gap only matters against a quality gap you have measured.
What if one run fails?
Read error on the failed receipt first. The documented codes tell you whether the cause is your input, a gated model or a platform fault. A failed run still carries whatever artifacts it made, and its usage shows what was spent, so record it as a data point rather than throwing it away: a model that fails often at a lower price is a real result.
If the POST itself fails with 400 invalid_request, check the model id against the catalog. Ids are the catalog strings, not the vendor's names, so claude-opus-5-5 will not work where openrouter/anthropic-claude-opus-5-5 does.
Keep the polling interval modest. The script sleeps 15 seconds, and a webhook is the documented alternative if you would rather not poll. Either way, stop on a terminal status and use expires_at as your ceiling, as the docs advise.
How many runs is enough?
There is no magic number, and the docs give none. A reasonable floor is three runs per model on two or three different inputs, so a single odd input does not decide the result. Agents are not deterministic, so two runs of one model on one input can differ in cost and in quality, and the spread tells you how much to trust the average.
Log the pair of numbers that matter, the debited cost and your quality score, in a spreadsheet with the receipt id and the model string. When the cheaper model is within your quality bar on most inputs, you have your answer. When it is not, you have learned that for this Format the extra spend buys something real.
Sources
Related posts
More in Developers
- Agent Completions 404 agent_run_not_found: which id to poll
A 404 agent_run_not_found means the id is not an Agent Completion run of your account, often an Action or Format run id. Poll, list and cancel agrun_ runs.
- Agent Completions 400: generation_spend_cap_usd is required
POST /v1/agent/completions returns 400 invalid_request without generation_spend_cap_usd. It has no default: why, how to size it, and what the receipt echoes.
- Agent Completions input_image: send a photo in messages[]
Send an image to Sume's Agent Completions as an input_image content part inside messages[], merge it with attachments, and get a caption back as typed output.
- AI API budget ran out: which error? Claude 400 and 429, Sume 402
Claude returns 429 enforced_spend_limit_reached at the tier cap and 400 for a limit you set. Sume returns 402 at create and a failed run at the cap.
Written by Sume