Check an Ideogram 4.5 edit left the rest of the image alone (Python)
Diff the source and the Ideogram 4.5 edit with Pillow, count changed pixels outside your text box, and fail the pass before it enters a chain. Code included.
To check that an Ideogram 4.5 edit left the rest of the image alone, download the source and the result as PNG, blank out the box where the text changed, and count the pixels that still differ. If more than a tiny share differ, reject the pass before it becomes the next source in a chain.
Ideogram launched 4.5 on 2026-09-30 saying it "eliminates artifact buildup, making multi-turn editing possible" (Ideogram on X, read 2026-10-05). That is a claim about the model, and this script is how you turn it into a measurement on your own pictures.
Why compare pixels when I can look at the picture?
Because you cannot see a 3 percent shift in a shadow, and a chain compounds it. A launch-partner page (Morphic, a partner and not the vendor, read 2026-10-05) says the model copies unchanged pixels from the source. Sume's docs make no such promise, so a diff is the only evidence that your result did.
What does the check look like?
The script takes two local files, a box you expect to change as left,top,right,bottom, and a tolerance per channel. An edit that omits aspect_ratio keeps the source geometry per the Image API docs, so the two sizes should match; if they do not, the script stops there, because a resize means every pixel moved.
import sys
from PIL import Image, ImageChops
def changed_outside(src, out, box, tol=8):
a = Image.open(src).convert("RGB")
b = Image.open(out).convert("RGB")
if a.size != b.size:
return None # geometry changed
diff = ImageChops.difference(a, b).convert("L")
mask = diff.point(lambda v: 255 if v > tol else 0)
mask.paste(0, box) # ignore the edited region
return sum(1 for v in mask.getdata() if v) / (a.width * a.height)
if __name__ == "__main__":
src, out = sys.argv[1], sys.argv[2]
box = tuple(int(x) for x in sys.argv[3].split(","))
share = changed_outside(src, out, box)
if share is None:
sys.exit("FAIL: size differs")
print(f"{share:.4%} of pixels changed outside the box")
sys.exit(1 if share > 0.01 else 0)
What threshold should fail a pass?
There is no number in the docs, so pick yours from a baseline. Run the script on three edits you judged clean by eye and three you judged bad, and set the line between them. The table gives starting points to test, not facts about the model.
| Share changed outside the box | Verdict | Action |
|---|---|---|
| Under 0.1% | Clean | Use as the next source |
| 0.1% to 1% | Look at it | Open the diff mask, then decide |
| Over 1% | Reject | Re-run the pass, or tighten the prompt |
What if everything differs a little?
Check the format first. If you saved the source as a JPEG and the result is a PNG, the whole picture differs by compression noise, and your tolerance of 8 may or may not hide it. Keep every pass as a PNG; the PNG between passes post explains the loss. Note that Ideogram 4.5 does not take output_format, so you cannot ask for PNG; convert on your side after download.
Also check that the box is right. If the headline moved down a line, the old position now differs and the new one differs too, so the box has to cover both. Widen it before you widen the tolerance.
Where does this sit in a chain?
Run it after every pass and before you submit the next one. A failed check costs one edit you already paid for, $0.0375 at low on Sume, and saves every later pass built on a damaged source. If your edit is a region-masked job on a model that supports mask_url, the same check for GPT Image 2.5 applies with a stricter box.
How do I run it across a whole batch?
Wrap the function in a loop over (source, result, box) tuples and write one line per file to a CSV: file name, share changed, verdict. Sort by share and open only the worst five masks; that is faster than reviewing every picture and catches the failures that matter.
Save the diff mask for any failure with mask.save("diff.png"). A mask that shows a clean rectangle around the headline is a pass with a loose box, and a mask that shows speckle across the whole image is usually a format problem: a JPEG source against a PNG result. Keep your thresholds in one place, so you can tighten them as you collect examples of good and bad edits.
Sources
Related posts
More in Developers
- Check one video model first: GET /v1/video-router/models/{id}
Before you spend on a clip, read one model's limits with GET /v1/video-router/models/{model_id}. It returns capabilities and price, and 404s on an unknown id.
- Reference video job failed on download: check input URLs first
A Sume reference-to-video job fails when an input URL is not public. Python that HEAD-checks references before POST /v1/videos, and how to read the error.
- CI test: assert Seedance 2.5 and Wan 3.0 still list 30 seconds
Fail a build when your 30-second video code outlives the catalog. A short Python test reads supported_durations from GET /v1/videos/models and checks 30.
- Retry, fix or stop: classify every Sume /v1/videos error in code
One Python function that maps each documented /v1/videos error code (400, 402, 404, 409, 429, 502) to retry, fix or stop, plus the two 409s that look alike.
Written by Sume