Blind A/B test of two image edits: reviewer sheet and key in Python

Pick between GPT Image 2.5 and Nano Banana 2 edits without bias: shuffle left and right, hide model names, save an answer key. Pillow script, 25 lines.

5 min readSume
All posts

Both vendors say their new image model keeps your references faithfully. OpenAI says Images 2.5 improves consistency across multiple edits of an uploaded photo (OpenAI Developer Community, read 2026-10-04), and Google describes up to 4 character images and 10 object images for consistency in Nano Banana 2 (Google AI for Developers, read 2026-10-04). A buyer's way through that is a blind test on their own photos.

Run the same edits on both

Call openai/gpt-image-2.5 and google/nano-banana-2 on the same source photo and prompt through the Sume Image API, and save each output as <photo>-gpt.png and <photo>-nb2.png. Note that both are paid calls, so ten photos on two models is twenty images.

Blind-test design, read 2026-10-04
StepRuleWhy
OrderShuffle which model is on the leftRemoves a left-side habit
LabelsShow A and B onlyHides model names
KeyWrite the real mapping to a CSV you do not shareLets you unblind after scoring
QuestionWhich keeps the person or product more faithful?One question per sheet

Build the sheets

Send the sheets to three or more reviewers, collect A or B for each photo, then open the key.

import csv
import random
from pathlib import Path
from PIL import Image, ImageDraw

photos = ["shoe", "mug", "coat"]
rows = []
for name in photos:
    pair = [("gpt", Image.open(f"{name}-gpt.png")), ("nb2", Image.open(f"{name}-nb2.png"))]
    random.shuffle(pair)
    w, h = 640, 640
    sheet = Image.new("RGB", (w * 2, h + 40), "white")
    d = ImageDraw.Draw(sheet)
    for i, (label, img) in enumerate(pair):
        sheet.paste(img.convert("RGB").resize((w, h)), (i * w, 40))
        d.text((i * w + 10, 10), "AB"[i], fill="black")
    sheet.save(f"sheet-{name}.png")
    rows.append([name, pair[0][0], pair[1][0]])
with open("answer-key.csv", "w", newline="") as f:
    csv.writer(f).writerows([["photo", "A", "B"], *rows])

Reading the result

With ten photos a 7 to 3 split is weak evidence and a 10 to 0 split is strong. Report counts, not a headline claim, and include the photos where reviewers could not tell.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume