Did the edit stay in its region? A pixel-diff check for GPT Image 2.5
After a GPT Image 2.5 edit, measure how much changed outside the area you meant to change. A Pillow script that diffs the result against the original.
To check that a GPT Image 2.5 edit changed only the area you meant, download the result, resize it to the original's size, black out the region you edited in both pictures, and measure what still differs. A mean difference near zero outside the box means the rest held; a high value means the model touched more than you asked.
Both OpenAI's Image prompting guide and fal's GPT Image 2.5 guide (read 2026-10-02) warn that repeated edits can still change details you meant to keep, and say to inspect each result. Sume's Image API page lists mask_url for GPT Image 2.5 edits but does not publish a pixel-exact guarantee, so a measurement is worth running.
Why measure instead of looking?
A slight shift in a background texture or a tone change in skin is easy to miss by eye and easy to catch with a number. The check is cheap: it runs locally and costs nothing on Sume. It is also useful as a gate in a batch, where you cannot look at every image.
What does the script do?
You give it the original file, the result URL and the box you meant to edit as left, top, right, bottom in the original's pixels. The script resizes the result to the original's size, masks the box out of the difference and reports the mean change per channel. Run pip install pillow requests. The threshold is yours to choose; start by running it on a result you consider good and a result you consider bad.
import io, sys, requests
from PIL import Image, ImageChops, ImageDraw, ImageStat
orig_path, result_url = sys.argv[1], sys.argv[2]
box = tuple(int(v) for v in sys.argv[3:7]) # left top right bottom
orig = Image.open(orig_path).convert("RGB")
res = Image.open(io.BytesIO(requests.get(result_url, timeout=60).content))
res = res.convert("RGB").resize(orig.size)
diff = ImageChops.difference(orig, res)
ImageDraw.Draw(diff).rectangle(box, fill=(0, 0, 0)) # ignore edited area
mean = ImageStat.Stat(diff).mean
print("mean change outside the box (0-255) R,G,B:",
[round(v, 2) for v in mean])
print("changed pixels bbox:", diff.getbbox())How do I read the output?
The table is a rule of thumb. The numbers depend on the output size, the resize filter and the image, so calibrate on your own samples.
| Mean change | Likely meaning | Action |
|---|---|---|
| Close to 0 | The rest held | Ship, after a visual check |
| Small but not zero | Resize, compression or tiny tone shift | Compare the two on screen |
| Large in a band | The edit spread past the box | Tighten the prompt or add a mask_url |
| Large everywhere | The model re-rendered the scene | Restate must-not-move and retry |
What are the limits of this check?
If the model changed the output's shape, the resize hides a crop, so request aspect_ratio: "auto" on edits, as the Image API docs advise. A pixel diff also cannot judge a face; use it with the identity constraint in the same-face prompt block and your own eyes. Sume does not run this comparison for you. A failed generation is not billed, but a completed one with a leaky edit is, so measure before you pay for a second pass.
How do I pick the box?
Draw it a little larger than the object you changed, because a soft edit blends past the outline. If you also send a mask_url, make the box cover the same area as the mask. Keep a small log of the measured value per image; a rising trend across a batch is a sign that your prompt is leaving too much room for the model.
Sources
- Image API
- Image prompting (OpenAI, read 2026-10-02)
- [How To Use GPT Image 2.5: Prompts & Workflows [2026] (fal, read 2026-10-02)](https://fal.ai/learn/tools/how-to-use-gpt-image-2-5)
Related posts
More in Developers
- Verify a Sume avatar video webhook in Python (HMAC SHA-256)
A Python verifier for avatar video webhooks: timestamp tolerance, rotation-safe comparison, and a hard refusal when the signing secret is empty.
- low_confidence_long_video: why video_frames warns past 90 seconds
Sume's video_frames returns the low_confidence_long_video warning when the source runs over 90 s. The job still succeeds; the hard cap is 300 s. What to do.
- Check has_audio first: video_inspect frames false before STT or detach
A free probe-only video_inspect tells you probe.has_audio before you reserve STT or run audio detach, so silent clips never hit the no-audio errors.
- video_inspect silence_split_seconds: sentence segments for captions
How silence_split_seconds (0.2 to 3) shapes Sume video-inspect sentence segments, the 0.5 s default in the repo, and turning segments into caption cues.
Written by Sume