FLUX 3 style boxes for Sume: draw a layout reference sheet in Python

Sume has no bounding-box input. Draw numbered boxes on a blank sheet from a 0-1000 grid, pass it as a reference, and name each box in the prompt.

6 min readSume
All posts

To get a bounding-box layout on Sume, draw it yourself: render numbered, colored boxes on a blank canvas, send that sheet as an input reference, and tell the prompt what each numbered box holds. Sume's image API has no box field, so this is a prompt technique, not a guarantee. It is the closest approach to what BFL describes for FLUX 3 Image, where boxes are given as [y0, x0, y1, x1] on a 0 to 1000 grid (read 2026-10-03).

Because the grid is relative, one layout works at any output size. The script below converts the grid to pixels, so you can use it for a 1536 x 1024 request or a 3840 x 2160 one.

From grid to sheet

Each element is a label and a box in BFL's order, y first. The function scales to the canvas, draws a colored outline with a number, and returns the text legend to paste into the prompt: Box 1, Box 2, and so on. The legend is what ties each rectangle to its content.

from PIL import Image, ImageDraw

def layout_sheet(size, elements, path="layout.png"):
    """Draw labelled boxes. Each element is (label, [y0, x0, y1, x1]) on a 0-1000 grid."""
    w, h = size
    im = Image.new("RGB", size, (255, 255, 255))
    d = ImageDraw.Draw(im)
    palette = [(230, 80, 80), (60, 130, 220), (70, 180, 110), (240, 170, 40)]
    for i, (label, (y0, x0, y1, x1)) in enumerate(elements):
        px = (x0 * w // 1000, y0 * h // 1000, x1 * w // 1000, y1 * h // 1000)
        d.rectangle(px, outline=palette[i % 4], width=6)
        d.text((px[0] + 12, px[1] + 12), f"{i + 1}: {label}", fill=palette[i % 4])
    im.save(path)
    return "\n".join(f"Box {i + 1}: {label}" for i, (label, _) in enumerate(elements))

if __name__ == "__main__":
    note = layout_sheet((1536, 1024), [
        ("a woman in a yellow coat", [380, 80, 960, 300]),
        ("a picnic blanket with fruit", [620, 330, 940, 760]),
        ("a golden retriever", [500, 780, 940, 960]),
    ])
    print(note)

Wording the prompt

Say what the sheet is and what to do with it, in this order: the reference is a layout guide, boxes mark where each subject goes, do not draw the boxes or numbers in the final image, then the legend and the scene. Without the instruction to omit the guide, models sometimes copy the colored outlines into the picture.

Use a model that accepts references. Per Sume's docs, ChatGPT Image 2.5 takes up to 16, other edit models up to 10, and Ideogram 4.5 up to five with the first being edited, which makes it a poor fit because the sheet would be the image it edits.

read 2026-10-03 from BFL's FLUX 3 Image docs and Sume's image docs
ConcernBFL FLUX 3 Image (vendor docs)Sume today
Box inputNative, 0-1000 grid [y0, x0, y1, x1]None; send a drawn sheet as a reference
Box semanticsNew, Anchor, Move element typesDescribe each in words
EnforcementModel is built to follow the boxesSoft guidance; check the result
Reference limitUp to 1010 on most edit models, 16 on ChatGPT Image 2.5

Scaling the sheet to your request

Make the sheet the same aspect ratio as the output you ask for, since a mismatch means the boxes are stretched when the model reads them. Keep the sheet itself modest, 1024 pixels on the long edge is plenty, because it is guidance rather than content. A reference that is uploaded must be a public HTTPS URL on Sume, so host the PNG before the request.

Keep box labels short and concrete: noun plus one adjective. Long labels crowd small boxes and invite the model to read the sheet as text content to reproduce.

Check the result

Measure placement instead of eyeballing: run a detector or just open the image with the boxes overlaid, and look at each subject's center against its box. If a subject lands outside its box more than a few percent of the canvas, tighten the wording or reduce the number of boxes. Four or five elements is a sensible ceiling for a soft guide; beyond that, placement tends to blur.

A look at what Sume does offer for regions is in bounding boxes in the prompt versus a masked edit, and the mask route is in the box-to-mask post. The ad-layout version is layout-sensitive ads. Parameters are in the Image API docs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume