A/B testing three image models on 50 prompts: what it costs on Sume
Fifty prompts on ChatGPT Image 2.5 medium, Seedream 5.0 Lite and Nano Banana 2.1 cost $0.825, $2.1875 and $5.00 on Sume. The full test and a script to run it.

Running 50 prompts through three image models costs $8.01 billed on Sume: $0.825 on ChatGPT Image 2.5 at medium, $2.1875 on Seedream 5.0 Lite and $5.00 on Nano Banana 2.1 at 1K. That is the full price of an A/B/C test before you move a pipeline to a new model, and it is less than most teams spend on the meeting that decides it.
The numbers assume one image per prompt per model, a 1:1 ratio and no references. Add n or references and the total moves, so recompute before you run.
The cost, model by model
Each row multiplies the billed per-image price by 50. The GPT row uses a 1024x1024 image with a short prompt; with a longer prompt, add about $0.0001 per image at list.
| Model and setting | Billed per image | 50 images |
|---|---|---|
| ChatGPT Image 2.5, medium | $0.0165 | $0.825 |
| Seedream 5.0 Lite | $0.04375 | $2.1875 |
| Nano Banana 2.1, 1K | $0.10 | $5.00 |
| All three | $0.1603 | $8.01 |
Design the test so it can tell you something
Choose prompts from real work, not from a gallery. Fifty is enough to see a pattern if you sort them into five groups of ten: product, portrait, text on a sign, a scene with several objects, and an edit of an existing image. Hold the ratio and seed fixed where the model lists them, and name your outputs <prompt-id>-<model>.png so a reviewer can line them up.
Have each reviewer pick a winner per prompt without seeing the model name. Count wins by group, not just overall. A model that wins on portraits and loses on text tells you to route by job, not to pick one model.
What each line of the bill is for
The $0.825 on 2.5 medium is the cheapest line, and it also scales best: at 500 prompts the same test costs $8.25 on 2.5 medium, $21.88 on Seedream and $50.00 on Nano Banana. A first pass on all three at 50 prompts is worth it; a second pass on the winner at 500 is cheap enough to run before you commit.
If one model wins only one group of prompts, route that group to it and keep the cheaper row for the rest. The routing layer is a dictionary from job type to model id, and the request body does not otherwise change.
The script
The script below reads prompts from a list, submits each to the three ids and prints the billed amount reported on each result. It stops at 3 prompts here; replace the list with your 50.
import os, requests
URL = "https://api.sume.com/v1/images"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
MODELS = [("openai/gpt-image-2.5", {"quality": "medium"}),
("bytedance-seed/seedream-5-lite", {}),
("google/nano-banana-2.1", {"resolution": "1K"})]
PROMPTS = ["A red kettle on a stove", "A foggy harbor at dawn",
"A wooden sign that reads OPEN"]
for p in PROMPTS:
for model, extra in MODELS:
body = {"model": model, "prompt": p, "aspect_ratio": "1:1", **extra}
r = requests.post(URL, json=body, headers=H, timeout=120)
j = r.json()
print(p[:20], model, r.status_code, j.get("usage", {}).get("cost"))Common ways the test goes wrong
Three mistakes are worth avoiding. Sending quality to Seedream or resolution to Seedream returns 400 unsupported_parameter, since neither lists the field; build the per-model extras as in the script rather than one shared body. Leaving quality off for GPT Image 2.5 lands on the default, high, and the first row of the test costs $0.0659 per image, not $0.0165. And running the three models at different ratios makes the result impossible to compare, because 2.5's price depends on the ratio.
Reading the results
A call that takes longer than 30 seconds returns a 202 job envelope instead of the image, so a heavier model can return a job id on a slow day. Poll the job or add a handler for it before you scale the script to 50 prompts. Failed jobs are not billed, so a retry costs only the image that finally succeeds.
Sum usage.cost across the run and compare it with the table above. A difference means a prompt used a different tier or ratio than you planned. Keep the three totals next to each other in a spreadsheet; the cheap model's cost per kept image, not its cost per image, is the number to compare.
Keep in mind that every figure here is a billed price from Sume's catalog on the date in the table caption. Prices and limits can change, so before a large run, read the endpoint record for the exact model id and compare it with your plan. A one-minute check costs nothing, and it is the only way to be sure the number in your budget is the number on the invoice.
Sources
Related posts
More in Use cases
- Ad creative testing workflow with an API: five calls to a winner
A repeatable ad creative test on Sume: research hooks, generate variants, cut, caption and count the spend per variant. Five documented calls with prices.
- Add a hand holding the product to a packshot with GPT Image 2.5
Turn a packshot into a hand-held photo: send the product photo and a hand or pose reference to openai/gpt-image-2.5. $0.0835 at high on Sume for one edit.
- How do I add a voiceover and music to a timelapse video?
Narrate a 45-second timelapse with one TTS job, one music track and one Timeline render: 3 cents, $0.125 and $0.10, about 33 cents on Sume.
- Add sunglasses or a hat to a person in an existing video (Omni edit)
Use Gemini Omni Flash 1.1 video_url edit on Sume to add an accessory to a person already in a clip. Prompt wording, a frame check list, and 4-second prices.
Written by Sume