Compare 4 image models in one labelled 2x2 sheet with cost per tile

Send one prompt to four Sume image models, paste the results into a 2x2 Pillow sheet, and print each model id with its billed usage.cost on its own tile.

5 min readSume
All posts

Call POST /v1/images once per model with the same prompt, fit each result into a 512 pixel tile with ImageOps.fit, and draw the model id and usage.cost on a black strip at the bottom of the tile. Four calls with the models below cost about $0.24 in total, so you can run the sheet for every new prompt without worrying about budget.

Why a sheet beats four browser tabs

When you choose a model for a product shot, the question is rarely which one is best. It is which one is good enough at the lowest price for this prompt. A sheet puts the answer on one screen, with the price printed on the image so the screenshot you paste into a ticket carries its own evidence.

The four models here are all text-to-image capable. Recraft V4 is text-only in the Sume catalog and returns WebP, which Pillow opens without a flag, so it can sit in the same loop.

The script

Each model takes its own default size. ImageOps.fit crops to the tile ratio instead of squashing, so a wide result is cropped rather than distorted. The label uses Pillow's built-in bitmap font, which needs no font file.

import os, io, requests
from PIL import Image
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def gen(**body):
    r = requests.post("https://api.sume.com/v1/images", json=body, headers=H, timeout=60)
    r.raise_for_status()
    if r.status_code == 202:
        raise SystemExit("queued, read /v1/jobs/{id}/result: " + r.text)
    return r.json()
def fetch(u):
    return Image.open(io.BytesIO(requests.get(u, timeout=60).content))
from PIL import ImageDraw, ImageOps
MODELS = ["bytedance-seed/seedream-4.5", "black-forest-labs/flux.2-pro",
          "google/nano-banana-2", "recraft/recraft-v4"]
PROMPT = "Matte black espresso machine on a marble counter, morning light"
sheet = Image.new("RGB", (1024, 1024), "white")
total = 0.0
for i, m in enumerate(MODELS):
    out = gen(model=m, prompt=PROMPT)
    tile = ImageOps.fit(fetch(out["data"][0]["url"]).convert("RGB"), (512, 512))
    d = ImageDraw.Draw(tile)
    d.rectangle((0, 472, 512, 512), fill="black")
    d.text((8, 486), m + "  $" + str(out["usage"]["cost"]), fill="white")
    sheet.paste(tile, ((i % 2) * 512, (i // 2) * 512))
    total += out["usage"]["cost"]
sheet.save("sheet.png")
print("sheet cost", round(total, 4))

What each tile should cost

Planning numbers from the repo catalog at default quality and 1K output. The printed usage.cost is what you are billed.

Four-model sheet, list times 1.25, catalog values read 2026-10-05
Model idList per imageBilled per image
bytedance-seed/seedream-4.5$0.04$0.05
black-forest-labs/flux.2-pro$0.03$0.0375
google/nano-banana-2$0.08$0.10
recraft/recraft-v4$0.04$0.05
Sheet total$0.19$0.2375

Make the comparison fair

  • Use the same prompt, and leave aspect_ratio off so every model uses its default, or set one ratio all four list.
  • Run each model more than once if the decision matters. One sample hides variance, and n lets you get several takes per call where the model allows it.
  • Do not compare a text-only model against an edit model on an edit task. Reference-image prompts need a model that accepts input_references.
  • Record the date. Catalog prices and model behavior change, and a sheet from last quarter is not evidence this quarter.

Next steps

Once the sheet shows a clear winner, encode the choice in code with the model filter, and lock it in with a prompt regression test so a later model swap shows up as a diff. If you only need variants from one model, the contact sheet from n variants is the cheaper pattern.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume