Image model bake-off: 20 prompts on 6 Sume models costs $6.125

A fair test sends the same 20 prompts to each model. On Sume, Grok, Qwen, Flux 2 Pro, Seedream 5 Lite, Ideogram 4.5 and Nano Banana 2.1 cost $6.125 in all.

5 min readSume
All posts

Running the same 20 prompts through six Sume image models costs $6.125 in total: $0.30625 per prompt across the six, times 20. The cheapest model in the set costs $0.50 for the 20 prompts and the dearest, Nano Banana 2.1 at 1K, costs $2.00.

The cost is small enough that the test design, not the budget, is the hard part. This page gives the cost table, then the rules that make a comparison fair.

What each model costs for 20 prompts

The test is cheap because five of the six models bill under ten cents. The one exception, Nano Banana 2.1 at 1K, bills exactly ten cents.

Prices are Sume billed amounts at each model's default settings, which is the provider list price times 1.25, per the Image API docs. Ideogram 4.5 is at its default medium quality and Nano Banana 2.1 at its 1K tier.

Billed cost of a 20-prompt bake-off, one image per prompt (read 2026-10-07)
ModelBilled per image20 prompts
Grok Imagine$0.025$0.50
Qwen Image$0.025$0.50
Flux 2 Pro$0.0375$0.75
Seedream 5.0 Lite$0.04375$0.875
Ideogram 4.5 (medium)$0.075$1.50
Nano Banana 2.1 (1K)$0.10$2.00
All six$0.30625$6.125

Rules that keep the comparison fair

  • Use identical prompts. Do not tune the wording per model during the first round.
  • Fix the aspect ratio to one value all six list, such as 1:1, and read each model's capability descriptor to confirm it.
  • Send one image per prompt. Grok Imagine takes only one per request in the catalog, so n: 1 is the common setting.
  • Hide the model names when you rate results, or have a second person label them.
  • Record usage.cost for each response. It should match the table.

What the test cannot tell you

Twenty prompts give you a view, not a statistic. With one sample per prompt and no seed support on Sume, a model can look good or bad by chance. If two models are close, add 20 more prompts before you decide.

Because every row costs only cents, a second round is cheap. Doubling the test to 40 prompts costs $12.25, which is less than the price of 47 ChatGPT Image 2 images (47 x $0.26375 = $12.40).

Scoring the results

Pick the score before you look at images. A simple scheme is three yes or no questions per image: did it follow the prompt, is the subject free of obvious defects, and would you ship it. Sum the yes answers per model across the 20 prompts, then divide the model's cost by the number of images you would ship. That gives a cost per usable image, which is the number that matters when a cheap model has a low hit rate.

Keep the raw scores. When a new model arrives, add it as a seventh column and run only that column, rather than repeating the whole test.

Edit and text checks need the right models

Run the six calls for each prompt in parallel so the whole bake-off finishes in about the time of the slowest model, and write each result to a file named by model and prompt number so nothing is overwritten.

If you also want to compare reference edits, check each model's input_references range first. Soul, Imagen 4 Fast and Ultra, Recraft V4 and Qwen Image Max are text-only on Sume, so they cannot join an edit round. For lettering, make a separate prompt set with exact strings and score them by spelling, not by overall look.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume