Move an object in a photo with AI: erase, then place with two masks

No Sume image model has a move operation. Move an object by erasing it with one mask and placing it with a second, using GPT Image 2.5 and Python.

6 min readSume
All posts

To move an object in a photo with Sume, run two masked edits: the first erases the object at its old position, the second places it at the new one. Sume's Image API has no "move" parameter, and only the ChatGPT Image 2.5 models take a mask_url, so both passes go to openai/gpt-image-2.5 or its Sunburst variant.

Black Forest Labs ships a Move element type in FLUX 3 Image, which takes a box at the source and a box at the target in one request. If one-request moves are your requirement, that is a reason to look at BFL; the recipe below is what you can run on Sume today.

Step 1: build both masks from boxes

FLUX-style boxes are [y0, x0, y1, x1] on a 0 to 1000 grid. Convert them to pixels for your image size, then draw a mask that is opaque everywhere and transparent inside the box. That is OpenAI's convention for edit masks. Sume's docs say only that mask_url is a public HTTPS URL and state no format rule, so test one mask on a throwaway image before you batch.

Host the two PNG files at public HTTPS URLs; localhost and private addresses are rejected before submission.

from PIL import Image

def box_to_px(box, size):
    """FLUX-style [y0, x0, y1, x1] on a 0-1000 grid -> (x0, y0, x1, y1) pixels."""
    y0, x0, y1, x1 = box
    w, h = size
    return (round(x0 * w / 1000), round(y0 * h / 1000), round(x1 * w / 1000), round(y1 * h / 1000))

def edit_mask(size, box, path):
    """RGBA mask: opaque everywhere, transparent inside the box (OpenAI's convention)."""
    mask = Image.new("RGBA", size, (0, 0, 0, 255))
    x0, y0, x1, y1 = box_to_px(box, size)
    mask.paste((0, 0, 0, 0), (x0, y0, x1, y1))
    mask.save(path)

if __name__ == "__main__":
    size = (1536, 1024)            # read it from your source image
    old_box = [520, 140, 820, 340]   # where the object is now
    new_box = [500, 640, 800, 840]   # where it should end up
    edit_mask(size, old_box, "mask_old.png")
    edit_mask(size, new_box, "mask_new.png")
    print(box_to_px(old_box, size), box_to_px(new_box, size))

Step 2: erase, then place

Pass one asks the model to fill the old box with the surrounding background. Pass two takes the pass-one result as image 1 and the original photo as image 2, so the object itself is still available as a reference, and asks for it to be placed in the new box with matching light. Numbering the references in the prompt follows Sume's guidance for multi-reference edits.

Use aspect_ratio: "auto" on edits so the output keeps the source shape; Sume's docs say omitting the field is not the same as auto. quality: "high" is the default on Sume when you omit it for this model, and it is worth stating when the object has fine detail.

# pass 1 erases the object, pass 2 places it at the new box (SOURCE_URL, MASK_*_URL are public HTTPS URLs)
import os
import requests

def generate(payload):
    r = requests.post(
        "https://api.sume.com/v1/images",
        json=payload,
        timeout=60,
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    )
    if r.status_code == 202:
        raise RuntimeError("still running: " + r.json()["data"]["status_url"])
    r.raise_for_status()
    return r.json()["data"][0]["url"]

ref = lambda url: {"type": "image_url", "image_url": {"url": url}}
base = {"model": "openai/gpt-image-2.5", "quality": "high", "aspect_ratio": "auto"}

erased = generate({**base, "input_references": [ref(os.environ["SOURCE_URL"])],
    "mask_url": os.environ["MASK_OLD_URL"],
    "prompt": "Remove the object inside the masked area and fill it with the surrounding floor and wall. Change nothing else."})
moved = generate({**base, "input_references": [ref(erased), ref(os.environ["SOURCE_URL"])],
    "mask_url": os.environ["MASK_NEW_URL"],
    "prompt": "Image 1 is the scene. Image 2 shows the object. Place that exact object inside the masked area, matching light and shadow. Change nothing else."})
print(moved)

What to check afterwards

Two generations means two chances to change pixels you did not mean to touch. Compare the result to the original outside both boxes before you accept it, and treat shadows as part of the object: a moved mug with its shadow left behind at the old spot is the most common failure. Widen the old-position mask enough to cover the shadow.

read 2026-10-03
CheckWhy it mattersFix if it fails
Old spot is cleanPass one left a ghost or shadowWiden the first mask and rerun pass one
Object matches the originalPass two redrew label or shapeAdd a close-up of the object as a third reference
Rest of the photo unchangedA mask is guidance, not a hard clipComposite your original back everywhere outside both boxes
Size and perspectiveObject looks pastedMake the new box match the scale and floor line

When this is the wrong tool

If you can cut the object out yourself, a Pillow paste is exact and free, and the model only has to fill the hole. Use the two-pass edit when lighting, reflections or occlusion mean a plain paste will look wrong. For the composite step, edit one region and composite the rest back shows the pixel-exact route, and why an edit can leak outside the mask explains why you should not skip it.

The request fields used here are described in the Image API docs; results above 30 seconds return a job to poll, covered in Jobs and results.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume