Blind A/B test of two image edits: reviewer sheet and key in Python
Pick between GPT Image 2.5 and Nano Banana 2 edits without bias: shuffle left and right, hide model names, save an answer key. Pillow script, 25 lines.

Both vendors say their new image model keeps your references faithfully. OpenAI says Images 2.5 improves consistency across multiple edits of an uploaded photo (OpenAI Developer Community, read 2026-10-04), and Google describes up to 4 character images and 10 object images for consistency in Nano Banana 2 (Google AI for Developers, read 2026-10-04). A buyer's way through that is a blind test on their own photos.
Run the same edits on both
Call openai/gpt-image-2.5 and google/nano-banana-2 on the same source photo and prompt through the Sume Image API, and save each output as <photo>-gpt.png and <photo>-nb2.png. Note that both are paid calls, so ten photos on two models is twenty images.
| Step | Rule | Why |
|---|---|---|
| Order | Shuffle which model is on the left | Removes a left-side habit |
| Labels | Show A and B only | Hides model names |
| Key | Write the real mapping to a CSV you do not share | Lets you unblind after scoring |
| Question | Which keeps the person or product more faithful? | One question per sheet |
Build the sheets
Send the sheets to three or more reviewers, collect A or B for each photo, then open the key.
import csv
import random
from pathlib import Path
from PIL import Image, ImageDraw
photos = ["shoe", "mug", "coat"]
rows = []
for name in photos:
pair = [("gpt", Image.open(f"{name}-gpt.png")), ("nb2", Image.open(f"{name}-nb2.png"))]
random.shuffle(pair)
w, h = 640, 640
sheet = Image.new("RGB", (w * 2, h + 40), "white")
d = ImageDraw.Draw(sheet)
for i, (label, img) in enumerate(pair):
sheet.paste(img.convert("RGB").resize((w, h)), (i * w, 40))
d.text((i * w + 10, 10), "AB"[i], fill="black")
sheet.save(f"sheet-{name}.png")
rows.append([name, pair[0][0], pair[1][0]])
with open("answer-key.csv", "w", newline="") as f:
csv.writer(f).writerows([["photo", "A", "B"], *rows])
Reading the result
With ten photos a 7 to 3 split is weak evidence and a 10 to 0 split is strong. Report counts, not a headline claim, and include the photos where reviewers could not tell.
Sources
Related posts
More in Developers
- Blind-test Sonic 3.6 against 3.5 on your own script
A vendor's blind-test percentage is not yours. Render the same lines with two catalog versions through the TTS Router, shuffle them, and let listeners vote.
- Bluesky avatar and banner: 1,000,000 bytes, PNG or JPEG
Bluesky's profile record caps avatar and banner at 1,000,000 bytes each, PNG or JPEG only. How to size a GPT Image 2.5 output and check the bytes in Pillow.
- Bluesky link card thumbnail: 1,000,000 bytes, any image type
The thumb on a Bluesky link card is optional, accepts image/* and is capped at 1,000,000 bytes. Make a 1280x720 preview with Sume and shrink it under the cap.
- Browser voice app that starts Sume jobs: keep the key on your server
Voice apps run in the browser over WebRTC, but Sume keys belong on a server. A route handler that holds the key, allowlists models, reuses idempotency keys.
Written by Sume