AI object remover from a photo: mask_url on GPT Image 2.5

To remove an object from a photo through the Sume API, send the photo and a public mask_url to POST /v1/images with a GPT Image 2.5 model, then prompt the fill.

4 min readSume
All posts

Sume has no one-click remover. You remove an object with an edit call: send POST /v1/images with a ChatGPT Image 2.5 model, the photo as an image reference, a public HTTPS mask_url that marks the object, and a prompt that says what should fill the gap.

Sume facts are from the Image API docs; the mask behavior of GPT Image is from OpenAI's image guide. Both were read 2026-09-30.

Which models accept a mask?

Only ChatGPT Image 2.5. The docs list openai/gpt-image-2.5 (Flare) and openai/gpt-image-2.5-sunburst as supporting up to 16 image references and an optional mask_url. The parameter is described as an optional public HTTPS mask URL for ChatGPT Image 2.5 edits.

Other models do not list mask_url. A request that sets a parameter the selected model does not list is rejected with 400 unsupported_parameter rather than silently dropped, so a wrong model fails loudly instead of ignoring your mask.

What does the mask have to look like?

Sume only requires a public HTTPS URL. The shape of the mask follows OpenAI's guide for GPT Image: the image and mask must be the same format and size (under 50MB), and the mask must contain an alpha channel. If you send several input images, the mask applies to the first.

The guide also says masking with GPT Image is entirely prompt-based: the model uses the mask as guidance but may not follow its exact shape with complete precision. Expect to check the edge around the removed object.

Mask rules, from the Sume docs and OpenAI's guide, read 2026-09-30
ItemRule
Models with mask_urlChatGPT Image 2.5 (Flare, Sunburst)
Mask locationPublic HTTPS URL
Mask and imageSame format and size, alpha channel on the mask
Multiple imagesMask applies to the first image
ExactnessPrompt-based guidance, not a pixel-exact cut

What does the request look like?

References go in input_references, as in the docs' image-to-image example. The prompt should describe the result, not only the deletion.

const response = await fetch("https://api.sume.com/v1/images", {
  method: "POST",
  headers: {
    Authorization: "Bearer " + process.env.SUME_API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "openai/gpt-image-2.5",
    prompt: "Remove the masked object and fill with the surrounding pavement",
    input_references: [
      { type: "image_url", image_url: { url: "https://example.com/photo.png" } },
    ],
    mask_url: "https://example.com/mask.png",
    aspect_ratio: "auto",
  }),
});
console.log(await response.json());

Why set aspect_ratio to auto?

On edit calls the docs say to prefer aspect_ratio: "auto" to match the reference, and that omitting the field is not the same as auto. Without it, the output can come back in a different shape than your photo and the mask no longer lines up with what you compare against.

What if I have no mask?

Describe the removal in the prompt alone, as a prompt-only edit with no mask_url. That works on any model that accepts references, but the model decides which pixels change. For a precise removal, a mask is the better tool. For a comparison of mask fields on the older route, see Image 1.0 mask_image_url vs mask_url.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume