Remove an object from a photo without a mask: Ideogram 4.5 prompt edit
Remove a bin, cable or stray person from a photo with no mask: send it to ideogram/ideogram-v4.5 and name the object. $0.0375 to $0.275 per image on Sume.

You can remove an object from a photo without drawing a mask by sending the photo to ideogram/ideogram-v4.5 as the only input_references entry and naming the object and its position in the prompt, for example: remove the green bin at the left edge and continue the fence behind it. On Sume that costs $0.0375 at low, $0.075 at medium or $0.275 at high per image. Ideogram lists object removal among the things its 4.5 edit model does.
The catch is that a prompt-only edit can change more than the object. If a mask is acceptable, openai/gpt-image-2.5 accepts mask_url and is the more surgical option. This post explains when the prompt-only route is enough and how to tell.
What Ideogram 4.5 on Sume accepts
With references, the model edits the first image and treats up to four more as references. It does not accept mask_url, output_format or seed; the catalog lists only the parameters a model supports, and anything else returns 400 unsupported_parameter instead of being silently dropped. An edit without aspect_ratio keeps the source geometry, so the frame of your photo is preserved.
That leaves the prompt as the only way to say where the object is. Use landmarks that anyone could find in the picture: the bin at the left edge, the cable crossing the wall above the door, the person in the red jacket by the lamp post.
Write the prompt as three clauses
A reliable pattern is: what to remove, what to continue into the gap, what stays. Remove the green bin at the left edge. Continue the wooden fence and the grass behind it. Keep everything else in the photo exactly as it is, including the lighting and the people.
Be specific about the fill. If you do not say what belongs behind the object, the model picks something, and it may be a second bin. If the object casts a shadow or has a reflection, name those too, since a removed bin with a shadow left on the pavement is a tell.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "ideogram/ideogram-v4.5",
"prompt": "Remove the green bin at the left edge and its shadow. Continue the wooden fence and grass behind it. Keep everything else unchanged.",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/garden.jpg"}}
],
"quality": "medium"
}'Prompt-only versus masked
The two routes differ in what they promise. A mask limits the model to a region; a prompt asks it to behave. On a simple background such as a lawn, sky or plain wall, the prompt-only edit is usually fine and costs nothing extra. On a busy background with text, faces or repeating patterns, the model may redraw nearby details.
| Route | Model id | Region control | Per image |
|---|---|---|---|
| Prompt only, low | ideogram/ideogram-v4.5 | Prompt | $0.0375 |
| Prompt only, medium | ideogram/ideogram-v4.5 | Prompt | $0.075 |
| Prompt only, high | ideogram/ideogram-v4.5 | Prompt | $0.275 |
| Masked edit, medium (1024x1024, one input) | openai/gpt-image-2.5 | mask_url | $0.021 |
| Masked edit, high (1024x1024, one input) | openai/gpt-image-2.5 | mask_url | $0.0835 |
Check what else changed
Diff the output against the source. A clean removal lights up only the object and its shadow. If signs, faces or textures far from the object also light up, the model redrew them, and that is your cue to switch to a mask. Sume bills a completed edit even when the result is bad, and does not bill a failed job, so a quick pilot at low on three images is a cheap way to learn how your photo type behaves.
Shadows, reflections and repeats
Most failed removals fail on the things that surround the object, not the object. A removed bin leaves a shadow on the pavement. A removed lamp leaves a reflection in a puddle. A removed person leaves a gap in a row of repeated items, like railings or windows, and the model may fill it with the wrong number of bars. Name these in the prompt: and the shadow it casts, and its reflection in the glass.
For repeated structures, describe the rhythm you want continued: evenly spaced vertical railings, same spacing as the rest. If you are removing something from the foreground of a landscape, mention the horizon line and the ground texture so the fill is continuous. You can also run the edit twice, once to remove and once to tidy, but start each from the original photo with a better prompt rather than chaining edits, because each pass on an already edited image is another chance to drift.
Rules of thumb
Small objects on simple backgrounds: prompt-only at low. Large objects or objects that touch other things, such as a person leaning on a wall: prompt-only at high or a masked edit. Anything legal, such as removing a competitor's sign from your own shop photo: keep the original and make sure the edited version is not presented as unedited where that matters.
Sources
Related posts
More in Models
- Replace one prop in a video with AI: bottle to apple edit prompt
Swap one object in a finished clip with gemini-omni-flash-1.1 on Sume: the docs example prompt, fields you cannot set, and an 8-second price.
- Square 1:1 ad video: which Sume models accept it, and 4:5
Seedance, Kling, Wan and MiniMax accept 1:1 on Sume; Gemini Omni takes only 16:9 and 9:16. No video model lists 4:5, so Feed needs a crop. Table and steps.
- MiniMax H3 Max on Sume: video, lip-sync and recast, which id to call
Three Sume ids carry the MiniMax H3 Max name: minimax-h3-max for video, a lip-sync route and h3-max-recast. What each takes, its window and its rate.
- Video edit prompt: say what stays, then what changes (Omni Flash)
A prompt pattern for Gemini Omni Flash 1.1 video edit on Sume: one change per pass, an explicit keep clause, and what the video_url route fixes for you.
Written by Sume