Compare AI image models on the same prompt: a 20-line API script

MAI-Image-2.6 and Muse Image both claim No. 2 on Arena. Skip the leaderboard: run one prompt through several Sume image models and compare the URLs and cost.

5 min readSume
All posts

To compare image models fairly, send one identical prompt, aspect ratio and reference set to each model id and compare the outputs and the billed cost. On Sume that is one POST /v1/images per model with only model changing. The response gives a hosted image URL in data[].url and the amount billed in usage.cost, so you can line up look and price in one table.

Leaderboard rank is a poor substitute for this. Microsoft says MAI-Image-2.6 ranks No. 2 on Arena for text-to-image (Microsoft AI, read 2026-10-01), and Meta says Muse Image ranked No. 2 on Arena as of July 5, 2026 (Meta AI, read 2026-10-01). Both can be true, and neither says which model draws your product label correctly.

What should I hold constant?

Change one thing at a time.

Comparison controls, from the Sume Image API docs
SettingHold it fixed because
promptThe wording changes what each model draws
aspect_ratioModels list different ratios; pick one all share
input_referencesSame public HTTPS URLs for every model that accepts them
nUse 1 so cost compares per image
qualityOnly on models that list it; omit it for a cross-vendor run

What does the script look like?

Models that reject a field return 400, so the script prints the status and continues. It uses three ids from Sume's list; swap in any id from GET /v1/images/models.

import os
import requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODELS = ["openai/gpt-image-2.5", "google/nano-banana-2",
          "black-forest-labs/flux.2-pro"]
PROMPT = "product photo of a glass perfume bottle labelled NOIR, studio light"

for model in MODELS:
    r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=60,
        json={"model": model, "prompt": PROMPT, "aspect_ratio": "1:1"})
    if r.status_code != 200:
        print(model, r.status_code)
        continue
    body = r.json()
    print(model, body["usage"]["cost"], body["data"][0]["url"])

How do I judge the results?

Score on what you ship: is the label spelled right, is the product shape intact, does the lighting match your brand. Run three to five prompts, not one, and keep the URLs in a sheet. Text in the image is the most common failure, so include a short word in every prompt and read it back.

Limits

One run is anecdote; models vary between generations because Sume's image API does not serve a seed. Each call bills, so a five-model, five-prompt grid costs 25 images. A 202 means a model passed the 30-second sync window, so poll its job rather than counting it as a failure. MAI-Image-2.6 and Muse Image are not in Sume's image list as of this post.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume