Move an object in a photo with AI: erase, then place with two masks
No Sume image model has a move operation. Move an object by erasing it with one mask and placing it with a second, using GPT Image 2.5 and Python.

To move an object in a photo with Sume, run two masked edits: the first erases the object at its old position, the second places it at the new one. Sume's Image API has no "move" parameter, and only the ChatGPT Image 2.5 models take a mask_url, so both passes go to openai/gpt-image-2.5 or its Sunburst variant.
Black Forest Labs ships a Move element type in FLUX 3 Image, which takes a box at the source and a box at the target in one request. If one-request moves are your requirement, that is a reason to look at BFL; the recipe below is what you can run on Sume today.
Step 1: build both masks from boxes
FLUX-style boxes are [y0, x0, y1, x1] on a 0 to 1000 grid. Convert them to pixels for your image size, then draw a mask that is opaque everywhere and transparent inside the box. That is OpenAI's convention for edit masks. Sume's docs say only that mask_url is a public HTTPS URL and state no format rule, so test one mask on a throwaway image before you batch.
Host the two PNG files at public HTTPS URLs; localhost and private addresses are rejected before submission.
from PIL import Image
def box_to_px(box, size):
"""FLUX-style [y0, x0, y1, x1] on a 0-1000 grid -> (x0, y0, x1, y1) pixels."""
y0, x0, y1, x1 = box
w, h = size
return (round(x0 * w / 1000), round(y0 * h / 1000), round(x1 * w / 1000), round(y1 * h / 1000))
def edit_mask(size, box, path):
"""RGBA mask: opaque everywhere, transparent inside the box (OpenAI's convention)."""
mask = Image.new("RGBA", size, (0, 0, 0, 255))
x0, y0, x1, y1 = box_to_px(box, size)
mask.paste((0, 0, 0, 0), (x0, y0, x1, y1))
mask.save(path)
if __name__ == "__main__":
size = (1536, 1024) # read it from your source image
old_box = [520, 140, 820, 340] # where the object is now
new_box = [500, 640, 800, 840] # where it should end up
edit_mask(size, old_box, "mask_old.png")
edit_mask(size, new_box, "mask_new.png")
print(box_to_px(old_box, size), box_to_px(new_box, size))Step 2: erase, then place
Pass one asks the model to fill the old box with the surrounding background. Pass two takes the pass-one result as image 1 and the original photo as image 2, so the object itself is still available as a reference, and asks for it to be placed in the new box with matching light. Numbering the references in the prompt follows Sume's guidance for multi-reference edits.
Use aspect_ratio: "auto" on edits so the output keeps the source shape; Sume's docs say omitting the field is not the same as auto. quality: "high" is the default on Sume when you omit it for this model, and it is worth stating when the object has fine detail.
# pass 1 erases the object, pass 2 places it at the new box (SOURCE_URL, MASK_*_URL are public HTTPS URLs)
import os
import requests
def generate(payload):
r = requests.post(
"https://api.sume.com/v1/images",
json=payload,
timeout=60,
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
)
if r.status_code == 202:
raise RuntimeError("still running: " + r.json()["data"]["status_url"])
r.raise_for_status()
return r.json()["data"][0]["url"]
ref = lambda url: {"type": "image_url", "image_url": {"url": url}}
base = {"model": "openai/gpt-image-2.5", "quality": "high", "aspect_ratio": "auto"}
erased = generate({**base, "input_references": [ref(os.environ["SOURCE_URL"])],
"mask_url": os.environ["MASK_OLD_URL"],
"prompt": "Remove the object inside the masked area and fill it with the surrounding floor and wall. Change nothing else."})
moved = generate({**base, "input_references": [ref(erased), ref(os.environ["SOURCE_URL"])],
"mask_url": os.environ["MASK_NEW_URL"],
"prompt": "Image 1 is the scene. Image 2 shows the object. Place that exact object inside the masked area, matching light and shadow. Change nothing else."})
print(moved)What to check afterwards
Two generations means two chances to change pixels you did not mean to touch. Compare the result to the original outside both boxes before you accept it, and treat shadows as part of the object: a moved mug with its shadow left behind at the old spot is the most common failure. Widen the old-position mask enough to cover the shadow.
| Check | Why it matters | Fix if it fails |
|---|---|---|
| Old spot is clean | Pass one left a ghost or shadow | Widen the first mask and rerun pass one |
| Object matches the original | Pass two redrew label or shape | Add a close-up of the object as a third reference |
| Rest of the photo unchanged | A mask is guidance, not a hard clip | Composite your original back everywhere outside both boxes |
| Size and perspective | Object looks pasted | Make the new box match the scale and floor line |
When this is the wrong tool
If you can cut the object out yourself, a Pillow paste is exact and free, and the model only has to fill the hole. Use the two-pass edit when lighting, reflections or occlusion mean a plain paste will look wrong. For the composite step, edit one region and composite the rest back shows the pixel-exact route, and why an edit can leak outside the mask explains why you should not skip it.
The request fields used here are described in the Image API docs; results above 30 seconds return a job to poll, covered in Jobs and results.
Sources
Related posts
More in Use cases
- Movember 2026 team update videos from phone photos, one per week
Movember asks people to talk about men's health. Make four weekly update clips for a fundraising team page from phone photos, with caption cues you wrote.
- Movie and TV clip Shorts: what YouTube's help calls original
Unedited movie or TV clips are ineligible on Shorts. What YouTube's pages allow instead, and how to cut trailer-style Shorts from your own film or series.
- Multi-speaker dub: 32 speakers on ElevenLabs, a voice per line on Sume
ElevenLabs dubbing handles up to 32 speakers per file. On Sume you give each speaker a voice, make TTS jobs per line, and join up to 20 parts per concat.
- Music bed longer than your Short: YouTube Create caps it, Sume cuts it
YouTube Create says audio cannot exceed the video's length. A Sume Timeline render ends at audio.duration_seconds, with a soundtrack fade up to 10 s.
Written by Sume