GPT Image 2.5 arena: run your own blind test through the API

Arena leaderboards rank images on other people's prompts. Here is a blind three-way test of Flare, Sunburst and GPT Image 2 on yours, with one script on Sume.

4 min readSume
All posts

This post quotes no arena rank, because a leaderboard score says how voters judged other people's prompts, not how the model does on yours. To settle GPT Image 2.5 for your own work, send one prompt to openai/gpt-image-2.5, openai/gpt-image-2.5-sunburst and openai/gpt-image-2 on Sume, shuffle the three results and judge them without labels.

Model ids, quality values and response shape come from Sume's Image API docs and catalog code, read 2026-09-29.

Which models go into the test?

Flare is openai/gpt-image-2.5 and Sunburst is openai/gpt-image-2.5-sunburst. The docs also say ChatGPT Image 2 remains selectable, which gives you last generation as a baseline. Leave sume/auto out: Sume picks the family and never discloses which one ran, so you could not label the result.

Arms of a blind test on Sume, read 2026-09-29.
Arm`model` value`quality` values it lists
Flareopenai/gpt-image-2.5auto, low, medium, high, xhigh, max
Sunburstopenai/gpt-image-2.5-sunburstauto, low, medium, high, xhigh, max
ChatGPT Image 2openai/gpt-image-2low, medium, high

How do I keep the test fair?

Hold everything but the model constant. Use one prompt, one aspect_ratio and the same quality. Only low, medium and high exist on all three arms, and an omitted quality defaults to high on the 2.5 ids, so set high explicitly. Run a spread of prompts, not one: a product shot, a poster with text, an edit with a reference. Then score each set blind.

What does the script look like?

It generates one image per arm, writes the answer key to key.txt in shuffled order, and prints only numbered URLs. Open the URLs, pick the one you prefer, and read the key afterwards.

PROMPT="A poster for a jazz night, bold serif headline, warm lighting"
for m in openai/gpt-image-2.5 openai/gpt-image-2.5-sunburst openai/gpt-image-2; do
  url=$(curl -s -X POST https://api.sume.com/v1/images \
    -H "Authorization: Bearer $SUME_API_KEY" \
    -H "Content-Type: application/json" \
    -d "{\"model\":\"$m\",\"prompt\":\"$PROMPT\",\"quality\":\"high\"}" \
    | jq -r '.data[0].url')
  echo "$m $url"
done | sort -R > key.txt
awk '{print NR": "$2}' key.txt

What if an arm prints no URL?

Sume caps a blocking wait at 30 seconds. A generation that outlasts it returns a 202 job envelope with no data[0].url, so that arm prints nothing. Poll GET /v1/jobs/{id}/status and fetch the result, or add mode: "async" to every arm. Each response also carries usage.cost, the USD amount billed, so you can put cost next to quality in your notes.

How many prompts make the test meaningful?

One prompt proves nothing, because any model can win a single draw. Collect ten to twenty prompts that look like your real work: product photos, posters with text, edits of a supplied picture and a few hard cases you already know break models. Run the script once per prompt and keep the keys.

For each prompt, note the arm you preferred, or record a tie. Count wins per arm at the end, and add the usage.cost figures so the winner carries its price. If two arms tie on quality, the cheaper or faster one is the better choice for you. A small private tally like this answers the question an arena score cannot: what works on your prompts.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume