Compare AI image models on the same prompt: a 20-line API script
MAI-Image-2.6 and Muse Image both claim No. 2 on Arena. Skip the leaderboard: run one prompt through several Sume image models and compare the URLs and cost.

To compare image models fairly, send one identical prompt, aspect ratio and reference set to each model id and compare the outputs and the billed cost. On Sume that is one POST /v1/images per model with only model changing. The response gives a hosted image URL in data[].url and the amount billed in usage.cost, so you can line up look and price in one table.
Leaderboard rank is a poor substitute for this. Microsoft says MAI-Image-2.6 ranks No. 2 on Arena for text-to-image (Microsoft AI, read 2026-10-01), and Meta says Muse Image ranked No. 2 on Arena as of July 5, 2026 (Meta AI, read 2026-10-01). Both can be true, and neither says which model draws your product label correctly.
What should I hold constant?
Change one thing at a time.
| Setting | Hold it fixed because |
|---|---|
prompt | The wording changes what each model draws |
aspect_ratio | Models list different ratios; pick one all share |
input_references | Same public HTTPS URLs for every model that accepts them |
n | Use 1 so cost compares per image |
quality | Only on models that list it; omit it for a cross-vendor run |
What does the script look like?
Models that reject a field return 400, so the script prints the status and continues. It uses three ids from Sume's list; swap in any id from GET /v1/images/models.
import os
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODELS = ["openai/gpt-image-2.5", "google/nano-banana-2",
"black-forest-labs/flux.2-pro"]
PROMPT = "product photo of a glass perfume bottle labelled NOIR, studio light"
for model in MODELS:
r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=60,
json={"model": model, "prompt": PROMPT, "aspect_ratio": "1:1"})
if r.status_code != 200:
print(model, r.status_code)
continue
body = r.json()
print(model, body["usage"]["cost"], body["data"][0]["url"])How do I judge the results?
Score on what you ship: is the label spelled right, is the product shape intact, does the lighting match your brand. Run three to five prompts, not one, and keep the URLs in a sheet. Text in the image is the most common failure, so include a short word in every prompt and read it back.
Limits
One run is anecdote; models vary between generations because Sume's image API does not serve a seed. Each call bills, so a five-model, five-prompt grid costs 25 images. A 202 means a model passed the 30-second sync window, so poll its job rather than counting it as a failure. MAI-Image-2.6 and Muse Image are not in Sume's image list as of this post.
Sources
Related posts
More in Developers
- Why concurrency_limit differs from the Sume plan table
In Sume generation_limits, concurrency_limit is the effective cap and limit_source says plan or admin_override. Size work from it, not plan_concurrency_limit.
- curl --retry on a POST: retry a Sume submit with one key
curl --retry also retries a POST, and it resends the same headers each time. Put an Idempotency-Key on a Sume submit first, then pick --retry-max-time.
- Decart lucy-latest vs a pinned Lucy model; Sume catalog ids
Decart's lucy-latest alias can move while legacy Lucy Clip costs $0.15 per second against $0.04 for Lucy 2.5. Why pin a model id, and how to do it on Sume.
- Detect new AI image models: diff Sume GET /v1/images/models
Image models arrive weekly. A short Python diff against GET /v1/images/models tells you when Sume adds or retires an image model id, with no news feed to watch.
Written by Sume