Add an object to a set spot in a photo: box to words in Python
Sume has no bounding-box field. Turn a 0-1000 box into placement words for an Ideogram 4.5 edit, and into pixels for a mask, with a short Python helper.

To put an object at a chosen spot in a photo through the Sume Image API, write the position in words and, if the spot must be exact, send a mask. The API has no bounding-box field. Black Forest Labs' Flux 3 Image does: TechTimes reports that it takes a JSON list of elements, each with coordinates on a 0-to-1000 grid measured from the top-left corner, in the order [top, left, bottom, right]. Sume's Image API docs list Flux 2 Pro as black-forest-labs/flux.2-pro and do not list Flux 3 Image, so a box cannot be sent as a box today.
What you can reuse is the box as a way to think. Draw it once in your own tool, keep the numbers, and convert them two ways: into a short phrase for the edit prompt, and into pixels for a mask or a check.
Grid to words and pixels
The helper below splits the frame into thirds, names the cell that holds the box centre, and reports the box as a share of the frame. It also gives the pixel box for a given image size. It uses only the standard library, so you can run it as is.
def place_words(box, width, height):
top, left, bottom, right = box # 0-1000 grid, [top, left, bottom, right]
cy, cx = (top + bottom) / 2, (left + right) / 2
row = ["upper", "middle", "lower"][min(int(cy // 334), 2)]
col = ["left", "centre", "right"][min(int(cx // 334), 2)]
area = round((bottom - top) * (right - left) / 10_000)
px = (round(left * width / 1000), round(top * height / 1000),
round(right * width / 1000), round(bottom * height / 1000))
text = (f"in the {row} {col} of the frame, about {area}% of the frame area, "
f"its bottom edge {round(bottom / 10)}% of the way down")
return text, px
text, px = place_words([650, 60, 950, 360], 1600, 1200)
print(text)
print("pixel box (left, top, right, bottom):", px)The request
For an edit on Ideogram 4.5, send the photo as the first entry of input_references and say where the new object goes. Per the Sume docs, with references the model edits the first image and uses up to four more as references. Omit aspect_ratio so the result keeps the source shape. Use the helper's sentence as the position clause, for example: "Add a potted fern in the lower left of the frame, about 9% of the frame area, keep everything else unchanged."
Send the fern photo as a second reference if you want that exact plant. Run it at quality: "low" first. Sume sells Ideogram 4.5 at the provider list times 1.25, which works out to $0.0375 for a low-quality image, and the endpoint's pricing line is the figure to trust.
What the words cannot promise
A phrase is a request, not a constraint. The model can place the fern a little higher or larger, and nothing in the Sume docs says it keeps pixels outside the area unchanged. Check the result against your box: crop the pixel box from the output and look at it. If the spot must be exact, build a mask from the pixel box with Pillow and use mask_url on openai/gpt-image-2.5, the model the docs name for masked edits. The docs do not say which mask colour marks the edit area, so run one cheap test render before a batch.
| Need | Flux 3 Image (reported) | Sume Image API today |
|---|---|---|
| Where the object goes | Box on a 0-1000 grid, [top, left, bottom, right] | Position words in the prompt |
| Strict region | Pixels outside the box left unchanged, per the vendor | mask_url on GPT Image 2.5 only |
| Same layout at any size | Grid is independent of pixel size | Convert to pixels yourself |
| Several objects | Element IDs with their own boxes | One edit call per object |
For the coordinate maths on a full canvas, see the 0-1000 to pixels walkthrough. For the mask route, see GPT Image 2.5 mask_url edits.
Sources
Related posts
More in Developers
- Alt text for a 30-image gallery in one Sume Agent Completion
One Agent Completion call can take up to 30 images and return an alts array under an object schema. Python stdlib script with the cap, poll and limits.
- API key scopes for Sume: which key can call which endpoint family?
Sume API keys carry fixed scopes: formats:write, actions:read, agent_completions:write, account:read. Which scope each route needs, and why old keys get a 403.
- Arabic speech to text API: Sume STT with language_code ar
Transcribe Arabic audio with Sume STT: send language_code ar, check the reported language, and review the text. $0.01 per audio minute.
- Avatar job tracking table: which Sume ids to store and why
Avatar work produces a handle, a job id, a preview id and a video id. A small SQL table that keeps them straight, plus the status fields to poll.
Written by Sume