FLUX 3 bounding box to a mask_url: Python region edit on Sume

FLUX 3 Image boxes use [top, left, bottom, right] on a 0-1000 grid. Convert one to an RGBA mask with Pillow and run the region edit on Sume's GPT Image 2.5.

5 min readSume
All posts

A FLUX 3 Image bounding box is four numbers, [top, left, bottom, right], measured from the top-left corner of the image on a 0 to 1000 grid in both directions. To reuse one on Sume, scale the four numbers by the image's width and height, draw that rectangle on a transparent RGBA PNG, host it, and pass it as mask_url on a ChatGPT Image 2.5 edit. That gives you a region edit from the same coordinates, with one difference: the model treats the mask as guidance, not a boundary.

The script below does the conversion and prints the mask file you need to host. Sume does not accept coordinates, so this is the bridge.

What does the FLUX 3 box format actually say?

BFL's FLUX 3 Image overview shows boxes in edit examples as "tgt_bbox": [250, 50, 850, 650], in [top, left, bottom, right] order on a 0-1000 grid. Each element row also carries an id, a from reference such as ref_image_0, a src_bbox, a tgt_bbox, and a kind of new, anchor, or move. Tech Times describes the grid as normalized from the top-left corner, so a box does not depend on the output resolution (read 2026-10-03).

That resolution independence is what makes the conversion simple: the same box scales to any image size by multiplying by size / 1000.

What does Sume's mask edit need?

Sume lists mask_url as a public HTTPS URL for ChatGPT Image 2.5 edits (openai/gpt-image-2.5, also openai/gpt-image-2.5-sunburst), alongside up to 16 input_references. OpenAI's guide adds the file rules: the mask has an alpha channel, it matches the image's size and format, it stays under 50 MB, and with several images the mask applies to the first. The guide text I read does not state which alpha value marks the edit zone, so the script uses fully transparent as the edit area, the usual convention, and a flag to invert it.

FLUX 3 box versus Sume mask edit, read 2026-10-03
PropertyFLUX 3 Image boxSume mask_url edit
InputJSON with [top, left, bottom, right]A PNG with an alpha channel
Coordinate system0-1000 grid from top-leftPixels of the image
ShapesRectangles per elementAny painted shape
Moves an elementYes, via src_bbox and tgt_bboxNo; describe it in the prompt
StrictnessReported to preserve other pixelsGuidance, may not follow the exact shape

The conversion in Python

Set box to your four numbers and SRC to the local copy of the image you are editing. The mask is the same size as the source and opaque everywhere except the box. Host mask.png and the image over HTTPS, then send the second request. The script requires Pillow.

import os, requests
from PIL import Image, ImageDraw

SRC = "photo.png"
box = [250, 50, 850, 650]  # top, left, bottom, right on a 0-1000 grid
EDIT_ALPHA = 0  # alpha value that marks the area to change
W, H = Image.open(SRC).size
top, left, bottom, right = box
mask = Image.new("RGBA", (W, H), (0, 0, 0, 255 - EDIT_ALPHA))
rect = (left * W / 1000, top * H / 1000, right * W / 1000, bottom * H / 1000)
ImageDraw.Draw(mask).rectangle(rect, fill=(0, 0, 0, EDIT_ALPHA))
mask.save("mask.png")

# After hosting photo.png and mask.png at public HTTPS URLs:
IMG, MASK = "https://example.com/photo.png", "https://example.com/mask.png"
r = requests.post(
    "https://api.sume.com/v1/images",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    json={
        "model": "openai/gpt-image-2.5",
        "prompt": "replace the object in the masked area with a red ceramic vase",
        "input_references": [{"type": "image_url", "image_url": {"url": IMG}}],
        "mask_url": MASK,
        "aspect_ratio": "auto",
    },
    timeout=90,
)
print(r.status_code, r.json())

How loose should the rectangle be?

Because the mask is guidance, a rectangle that hugs the object tightly invites the model to leave the old edge visible, and a rectangle with generous padding gives it room to blend shadows and reflections. A margin of a few percent on each side is a reasonable first try: add 20 to the box's grid numbers and clamp to 0 and 1000 before you scale. Whatever you choose, keep the prompt specific about what belongs in the box and what must stay.

If your source is a box list produced by a layout planner, keep the original grid numbers and convert at the last step. That way the same list can be re-scaled for a different crop or aspect ratio, which is the reason BFL's format is resolution independent in the first place.

What does a failed edit tell you?

Sume bills a generation only when it completes, so a refused or failed edit costs nothing; the response is an error or a failed job, not a degraded image. If you get a 202, the work is still running: poll GET /v1/jobs/{id}/status and then /result, as the jobs docs describe. A 400 on the request usually means a field the model does not list, such as mask_url sent to a model other than ChatGPT Image 2.5.

What still differs after the conversion?

Three things. The box is one rectangle per element, so for several edits you either paint several rectangles into one mask and describe each in the prompt, or run the edits in sequence. A move, which BFL expresses with a source and target box, becomes two steps on Sume: erase at the source, then add at the target. And you lose the pixel-preservation claim that press coverage attaches to FLUX 3 Image; to get it back, paste the original over the result outside the mask, as in the composite-back recipe.

The bounding box versus masked edit note compares the two inputs in prose, and the mask checklist covers size and alpha preflight. If the wrong region changes, flip EDIT_ALPHA to 255 and rerun a low-quality test before you spend on a high-quality call.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume