GPT Image mask edit changed pixels outside the mask: what OpenAI says

OpenAI says GPT Image masks are guidance and may not follow the exact shape. Why a mask_url edit leaks, and how to keep the rest of the image on Sume.

5 min readSume
All posts

A GPT Image mask edit can change pixels outside the masked area because, per OpenAI, masking with GPT Image is entirely prompt-based: the model uses the mask as guidance but may not follow its exact shape. To keep the rest of the picture, say what must stay the same in the prompt, and if you need exact pixels, paste the edited region back over your original yourself.

What OpenAI and the press say

The OpenAI image generation guide says masking with GPT Image is entirely prompt-based and that the model may not follow the mask's exact shape with complete precision. It also says the image and mask must be the same format and size, and that masks need an alpha channel.

For the ChatGPT side, Winbuzzer's report on the Images 2.5 launch says highlighted areas are approximate and edits can extend beyond them. The same caveat applies to the API.

What OpenAI says about GPT Image masks (read 2026-10-02)
PointStatement
How the mask worksPrompt-based; used as guidance.
Shape accuracyMay not follow the exact shape with complete precision.
FormatImage and mask must be the same format and size.
AlphaMasks require an alpha channel.

What Sume adds

On Sume, GPT Image 2.5 takes an optional mask_url, a public HTTPS URL, on POST /v1/images. Sume's docs describe it as an optional mask for edits and do not promise pixel-exact results either. Send the original as an input reference, set aspect_ratio to auto so the frame matches, and write the preserve list into the prompt. For the mask rules, see the alpha-channel post and which image is masked when you send several.

Make the unchanged part exact

If the edit must leave everything outside a region untouched, composite. Keep your original, take the edited image, and paste the edited pixels back only where your own mask says. The snippet below uses Pillow, assumes the three files are the same size, and treats white in your own mask as the region to take from the edit.

from PIL import Image

original = Image.open("original.png").convert("RGB")
edited = Image.open("edited.png").convert("RGB")
region = Image.open("region_mask.png").convert("L")

for name, img in (("edited", edited), ("mask", region)):
    if img.size != original.size:
        raise SystemExit(f"{name} size {img.size} != {original.size}")

# white in region_mask.png = take the edited pixels
result = Image.composite(edited, original, region)
result.save("composited.png")
print("saved composited.png")

Limits to plan for

A composite fixes pixels, not lighting: a new object may need a shadow outside the mask, and a hard paste edge can show. Soften the edge of your own mask with a small blur and review the seam.

Billing stays per completed image, so a re-run to tighten the prompt costs a full image. Failed generations are not billed, per Sume's Image API docs.

Sources

Related posts

More in Models

All Models posts

Written by Sume