Compare orchestrator models in one Sume bulk queue with per-item model

A bulk item can carry its own model. Send the same input under several orchestrators, then compare debited spend and output across the receipts.

3 min readSume
All posts

Can one bulk queue test several models on the same input?

Yes. Each item in a bulk queue has the same body as a single run, and that includes model. It is validated when the queue is created, against the Agents catalog: an id outside the catalog is a 400 for the whole request, so you learn about a typo before any run starts.

Omit model and an item uses the default, which the call docs name as gpt-6-sol. The receipt shows the id that actually ran.

What does `model` change, and what does it not?

It selects only the orchestrator, the LLM that runs the agent turn: planning, tool calls, and the written answer. The Format's own tools pick the image, video and audio models. So an A/B across orchestrators compares how well each plans and follows the recipe, not how good the video model is.

That also explains the cost reading. The video generation is largely the same across items, while the agent turn differs, and the agent turn is exactly what usage.billable_amount_usd_micros leaves out. Compare usage.debited_usd_micros instead.

ids = ["MODEL_ID_A", "MODEL_ID_B", "MODEL_ID_C"]  # from the Agents catalog
same_input = {"product": "ceramic mug", "tone": "warm"}

items = [
    {"input": same_input, "model": m,
     "idempotency_key": f"ab-mug-{m}"}
    for m in ids
]
body = {"concurrency": len(items), "items": items}
print(len(body["items"]), "items")

How do I read the result fairly?

Per-item idempotency_key is accepted too, which is why the sample keys each item by model. Treat three items as an anecdote. Generated media varies run to run even with the same model, so repeat each model a few times before you declare a winner.

  • Keep the input identical across items and change only model.
  • Use debited_usd_micros for cost and the receipt's model id to confirm what ran.
  • Run the comparison on two or three different inputs, because one product can favor one orchestrator by accident.
  • Judge output by your own rubric, such as caption length and brand terms.
  • Remember queue-level idempotency hashes the whole items list, so changing one model makes a new payload.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume