Did the AI edit touch pixels outside the mask? Check it in numpy
BFL promises edits that leave the rest unchanged. Verify that claim on any model's output with a numpy diff outside your edit box, in about 15 lines.

To check whether an AI edit changed anything outside the area you meant to change, load the before and after images as arrays, compute the per-pixel difference, blank out your edit box, and look at what share of the rest differs. A result of 0.000 percent means the untouched region is identical; anything higher is drift you can measure instead of eyeball.
This matters now because FLUX 3 Image's launch material leans on preservation: BFL's docs describe an Anchor element type that is kept unchanged from a reference, and its multi-step edits are meant to leave the rest of the image alone. The same script works on any model's output, including the ones Sume serves, so you can test a claim on your own images rather than on a launch demo.
The check
The function below takes the original, the edited file and the edit box in pixels, and returns the share of outside pixels whose largest channel difference exceeds a tolerance. It raises if the output size changed, which is the first thing to rule out: an edit that returns a different canvas has resampled everything. The __main__ block builds a synthetic pair with one intended edit and one stray pixel so you can see both outcomes.
import numpy as np
from PIL import Image
def changed_outside(before_path, after_path, box_px, tol=0):
"""Share of pixels outside box_px (x0, y0, x1, y1) that differ by more than tol."""
a = np.asarray(Image.open(before_path).convert("RGB"), dtype=np.int16)
b = np.asarray(Image.open(after_path).convert("RGB"), dtype=np.int16)
if a.shape != b.shape:
raise ValueError(f"size changed: {a.shape} -> {b.shape}")
diff = np.abs(a - b).max(axis=2) > tol
x0, y0, x1, y1 = box_px
outside = np.ones(diff.shape, dtype=bool)
outside[y0:y1, x0:x1] = False
return diff[outside].mean()
if __name__ == "__main__":
base = Image.new("RGB", (400, 300), (200, 180, 160))
base.save("before.png")
edit = base.copy()
edit.paste((10, 10, 10), (100, 100, 200, 180)) # the intended edit
edit.putpixel((5, 5), (201, 180, 160)) # one stray pixel
edit.save("after.png")
print(f"{changed_outside('before.png', 'after.png', (100, 100, 200, 180)):.6%}")
print(f"{changed_outside('before.png', 'after.png', (100, 100, 200, 180), tol=2):.6%}")Reading the numbers
With tolerance 0, that example reports a tiny nonzero share: the single stray pixel I planted. With a tolerance of 2 levels it reports 0.000000 percent, because the stray pixel differs by only 1 level out of 255. Both answers are right; they ask different questions. Tolerance 0 asks "is it bit-identical", which is the claim a vendor makes when it says pixels are locked. A small tolerance asks "would a person notice", which is the question a client asks.
| Result outside the box | What it usually means | What to do |
|---|---|---|
| Exactly 0% | Region was left alone, or you composited it back | Ship it |
| Tiny % at tolerance 0, 0% at tolerance 2-3 | Re-encoding or color rounding | Fine for display; not bit-identical |
| A few % at tolerance 10+ | The model repainted soft areas near the edit | Composite the original back outside the box |
| Size or shape error | Output canvas differs from the input | Request aspect_ratio: "auto" on edits and resize check |
Three traps
Format first. Compare PNG to PNG. A JPEG or WebP result will fail a tolerance-0 test by construction, so ask for png where the model lists it as an output format; Sume's Image API accepts png, jpeg or webp on models that list formats, and a few rows (Ideogram 4.5, Higgsfield Soul) list none.
Second, edits that return a different size. Sume's docs say that on edit and image-to-image calls you should prefer aspect_ratio: "auto" to match the reference, and that omitting the field is not the same as auto. If the output is a different size, resize to the original with a high-quality filter before you diff, and expect drift everywhere.
Third, the box. Take your mask edges and add a margin of a few pixels, because models often blend a little past the edge of what you asked for. If you are checking a multi-element edit, build the outside region from the union of all the boxes.
What to do with the answer
If the check fails, the fix that does not depend on any vendor is to composite the original back everywhere outside the box. Edit one region and composite the rest back shows that step, and why masks are guidance explains why you cannot rely on a mask alone.
To compare models fairly, run the same edit and the same box through each one and keep the outside-drift number next to the cost. The edit fields and the models that take a mask are listed in the Image API docs.
Sources
Related posts
More in Developers
- Is my avatar ready? GET /avatars with status=ready and a handle filter
Check whether one Sume avatar handle is ready before an avatar video render, using the list route's status=ready and handle query parameters in Python.
- Chinese text to speech API: set language zh or it reads as English
Sume TTS only guesses Korean and Japanese when the language is missing. For Mandarin send language zh and pick a voice tagged zh, then test one line.
- Choose an AI video model in Python: filter the Sume model list by need
Read GET /v1/videos/models and keep only the models that fit your clip length, resolution and reference needs. A runnable Python filter for Sume.
- Claude Agent SDK 0.2.160: background wait ceiling, Sume video job
Claude Agent SDK Python 0.2.160 keeps stdin open until background subagents go idle, ten minutes by default. Why a Sume video job can outlast that.
Written by Sume