Flare vs Sunburst A/B test: 100 prompts cost $3.30 at medium
A 100-prompt A/B test of GPT Image 2.5 Flare and Sunburst costs $3.30 at medium on Sume, $13.18 at high. Budget by tier, plus a cost-logging script.

A fair A/B test of GPT Image 2.5's two models, Flare and Sunburst, over 100 prompts is 200 calls, and on Sume it costs $3.30 at medium and $13.18 at high, at 1024x1024. Both models quote the same price per image on Sume, so the budget depends only on the quality tier and the size you choose, not on which side wins.
OpenAI's changelog (read 2026-10-05) says the two models were released on September 8, 2026, with Sunburst prioritizing editing precision and Flare focusing on speedy, high-quality everyday images. Its deprecations page names either one as the replacement for gpt-image-1, which is why teams are testing both before the October 23 shutdown.
The budget
Each cell is 100 prompts times two models times Sume's price for that tier. The totals assume a fixed 1024x1024 size and a short prompt; long prompts add input text tokens at $5 per million.
| Quality | Price per image | One model, 100 prompts | Both models, 100 prompts |
|---|---|---|---|
| low | $0.0074 | $0.74 | $1.48 |
| medium | $0.0165 | $1.65 | $3.30 |
| high | $0.0659 | $6.59 | $13.18 |
| xhigh | $0.1171 | $11.71 | $23.43 |
| max | $0.2635 | $26.35 | $52.70 |
Run both sides with one script
The script sends each prompt to both model ids and totals usage.cost from the responses. Replace the two sample prompts with your own list, and read the key from the environment as shown so it never lands in the file. It handles the 202 envelope that Sume returns when a render takes longer than its 30-second wait.
import os
import requests
URL = "https://api.sume.com/v1/images"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODELS = ["openai/gpt-image-2.5", "openai/gpt-image-2.5-sunburst"]
PROMPTS = [
"A ceramic mug on a pale oak table, soft window light",
"A red bicycle leaning on a white brick wall",
]
total = 0.0
for prompt in PROMPTS:
for model in MODELS:
body = {"model": model, "prompt": prompt, "quality": "medium",
"image_size": "1024x1024"}
reply = requests.post(URL, headers=HEADERS, json=body, timeout=60)
reply.raise_for_status()
payload = reply.json()
if reply.status_code == 202:
print(model, "still rendering:", payload["data"]["job"]["id"])
continue
total += payload["usage"]["cost"]
print(model, payload["data"][0]["url"])
print(f"spent ${total:.4f}")Keep the test fair
Three controls matter more than the sample size.
- Pin
qualityandimage_sizeon both sides. Omittingqualitymeanshighon Sume, and omitting the size reserves the upper bound, so a missing field turns the cheap test into an expensive one. - Do not set
seed. Sume's Image API lists the field in its schema but no model advertises it, so a request that sends one returns400 unsupported_parameter, and you cannot make two runs identical. Compare many prompts instead of one. - Judge blind. Shuffle the outputs, hide the model id, and score them before you look at which side made them.
Reading the result
Flare is the model Sume's auto routing uses, so a tie is a reason to keep it. A clear win for Sunburst on edit prompts is a reason to send edits to Sunburst and text-to-image to Flare, since the price is the same either way.
Run the edit half separately with the same reference URL on both sides. An edit adds input image tokens at $8 per million, so an edit test costs more than the table above. Use at least ten reference images that differ in subject, lighting and background, because a model that edits one portrait well can still fail on a product shot or a cluttered room, and a single reference hides that.
After the test
Write the winning id into one config value, not into each call site. The test is cheap enough to repeat when OpenAI changes either model, and a single value lets you repeat it by changing one line.
If you only want the system to choose, send model: "sume/auto". Sume's Image API page says auto routing continues to use Flare and that the response never discloses which family ran, so use explicit ids for any test where the winner matters.
Save the raw outputs and the prompt list next to the score sheet. The URLs Sume returns are Sume-hosted, so download the files you want to keep before you close the test, then rerun the same list against the other model id whenever a new snapshot appears.
Sources
Related posts
More in Developers
- FLUX 3 Image 503: read the JSON status before you retry
BFL says a FLUX 3 Image 503 may be retryable once you check the JSON status; 422, 400 and 402 are not. Sume sync failures return a 502 with a retryable flag.
- FLUX 3 Image 0-1000 boxes to pixels on a 1920x1080 canvas
FLUX 3 Image boxes are [top, left, bottom, right] on a 0-1000 grid. The BFL example [250, 50, 850, 650] becomes x 96, y 270, 1152x648 pixels on 1920x1080.
- FLUX 3 Image images field: 2-10 references vs Sume's input_references
BFL's images field takes 2-10 references addressed as <ref_image_N>. Sume's input_references takes up to 10, or 16 on ChatGPT Image 2.5.
- Forgot the TTS language field? Sume only infers Korean and Japanese
Omit language on a Sume TTS request and the voice check assumes English. Sume infers ko or ja only when Hangul or kana outnumber Latin letters.
Written by Sume