Compare orchestrator models in one Sume bulk queue with per-item model
A bulk item can carry its own model. Send the same input under several orchestrators, then compare debited spend and output across the receipts.

Can one bulk queue test several models on the same input?
Yes. Each item in a bulk queue has the same body as a single run, and that includes model. It is validated when the queue is created, against the Agents catalog: an id outside the catalog is a 400 for the whole request, so you learn about a typo before any run starts.
Omit model and an item uses the default, which the call docs name as gpt-6-sol. The receipt shows the id that actually ran.
What does `model` change, and what does it not?
It selects only the orchestrator, the LLM that runs the agent turn: planning, tool calls, and the written answer. The Format's own tools pick the image, video and audio models. So an A/B across orchestrators compares how well each plans and follows the recipe, not how good the video model is.
That also explains the cost reading. The video generation is largely the same across items, while the agent turn differs, and the agent turn is exactly what usage.billable_amount_usd_micros leaves out. Compare usage.debited_usd_micros instead.
ids = ["MODEL_ID_A", "MODEL_ID_B", "MODEL_ID_C"] # from the Agents catalog
same_input = {"product": "ceramic mug", "tone": "warm"}
items = [
{"input": same_input, "model": m,
"idempotency_key": f"ab-mug-{m}"}
for m in ids
]
body = {"concurrency": len(items), "items": items}
print(len(body["items"]), "items")
How do I read the result fairly?
Per-item idempotency_key is accepted too, which is why the sample keys each item by model. Treat three items as an anecdote. Generated media varies run to run even with the same model, so repeat each model a few times before you declare a winner.
- Keep the input identical across items and change only
model. - Use
debited_usd_microsfor cost and the receipt's model id to confirm what ran. - Run the comparison on two or three different inputs, because one product can favor one orchestrator by accident.
- Judge output by your own rubric, such as caption length and brand terms.
- Remember queue-level idempotency hashes the whole
itemslist, so changing one model makes a new payload.
Sources
Related posts
More in Formats
- Cyber Monday email header: a magazine cover Format with room for type
sume-magazine-cover-campaign returns a cover-style still with space for headline text. Add the sale line in your email builder so it stays editable.
- Three Demand Gen video hooks in one Sume bulk run
Queue three hook variants of one video in a single Sume Format bulk run (up to 100 items per request) so you can test which opening works.
- Detect that someone edited your Sume Format before the weekly run
Pin the Format's package_sha and version, and compare them right before a weekly run. If a teammate edited the Format, stop the run instead of paying for it.
- Edit a Format mid-batch: which version do queued items use?
A bulk queue dispatches items over time. If you edit the Format while it runs, later items can run the new package. Here is how to tell from the receipts.
Written by Sume