Add an object to a set spot in a photo: box to words in Python

Sume has no bounding-box field. Turn a 0-1000 box into placement words for an Ideogram 4.5 edit, and into pixels for a mask, with a short Python helper.

5 min readSume
All posts

To put an object at a chosen spot in a photo through the Sume Image API, write the position in words and, if the spot must be exact, send a mask. The API has no bounding-box field. Black Forest Labs' Flux 3 Image does: TechTimes reports that it takes a JSON list of elements, each with coordinates on a 0-to-1000 grid measured from the top-left corner, in the order [top, left, bottom, right]. Sume's Image API docs list Flux 2 Pro as black-forest-labs/flux.2-pro and do not list Flux 3 Image, so a box cannot be sent as a box today.

What you can reuse is the box as a way to think. Draw it once in your own tool, keep the numbers, and convert them two ways: into a short phrase for the edit prompt, and into pixels for a mask or a check.

Grid to words and pixels

The helper below splits the frame into thirds, names the cell that holds the box centre, and reports the box as a share of the frame. It also gives the pixel box for a given image size. It uses only the standard library, so you can run it as is.

def place_words(box, width, height):
    top, left, bottom, right = box  # 0-1000 grid, [top, left, bottom, right]
    cy, cx = (top + bottom) / 2, (left + right) / 2
    row = ["upper", "middle", "lower"][min(int(cy // 334), 2)]
    col = ["left", "centre", "right"][min(int(cx // 334), 2)]
    area = round((bottom - top) * (right - left) / 10_000)
    px = (round(left * width / 1000), round(top * height / 1000),
          round(right * width / 1000), round(bottom * height / 1000))
    text = (f"in the {row} {col} of the frame, about {area}% of the frame area, "
            f"its bottom edge {round(bottom / 10)}% of the way down")
    return text, px

text, px = place_words([650, 60, 950, 360], 1600, 1200)
print(text)
print("pixel box (left, top, right, bottom):", px)

The request

For an edit on Ideogram 4.5, send the photo as the first entry of input_references and say where the new object goes. Per the Sume docs, with references the model edits the first image and uses up to four more as references. Omit aspect_ratio so the result keeps the source shape. Use the helper's sentence as the position clause, for example: "Add a potted fern in the lower left of the frame, about 9% of the frame area, keep everything else unchanged."

Send the fern photo as a second reference if you want that exact plant. Run it at quality: "low" first. Sume sells Ideogram 4.5 at the provider list times 1.25, which works out to $0.0375 for a low-quality image, and the endpoint's pricing line is the figure to trust.

What the words cannot promise

A phrase is a request, not a constraint. The model can place the fern a little higher or larger, and nothing in the Sume docs says it keeps pixels outside the area unchanged. Check the result against your box: crop the pixel box from the output and look at it. If the spot must be exact, build a mask from the pixel box with Pillow and use mask_url on openai/gpt-image-2.5, the model the docs name for masked edits. The docs do not say which mask colour marks the edit area, so run one cheap test render before a batch.

Box-based placement: Flux 3 Image as reported and what Sume offers (read 2026-10-07)
NeedFlux 3 Image (reported)Sume Image API today
Where the object goesBox on a 0-1000 grid, [top, left, bottom, right]Position words in the prompt
Strict regionPixels outside the box left unchanged, per the vendormask_url on GPT Image 2.5 only
Same layout at any sizeGrid is independent of pixel sizeConvert to pixels yourself
Several objectsElement IDs with their own boxesOne edit call per object

For the coordinate maths on a full canvas, see the 0-1000 to pixels walkthrough. For the mask route, see GPT Image 2.5 mask_url edits.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume