GPT Image 2.5 layout-sensitive ads: give it a layout reference image

OpenAI says GPT Image 2.5 can struggle with composition in layout-sensitive designs. On Sume, pass a wireframe as image 1 and fix one region with a mask.

5 min readSume
All posts

If a GPT Image 2.5 ad keeps putting the headline, product and logo in the wrong places, stop describing the layout in words and send a layout image. Sume accepts up to 16 reference images for ChatGPT Image 2.5, so a grey-box wireframe can be image 1, the product photo image 2, and the prompt can say which box holds what.

The reason to do this comes from OpenAI itself. Its image generation guide, read on 2026-10-03, lists known limitations for the GPT Image models, including difficulty with composition control in layout-sensitive designs, and says that text placement and clarity can still struggle even though the model is significantly improved. Sume's side is from the Image API docs.

What exactly does OpenAI say about layout and text?

Two statements matter. First, the guide says that although text rendering is significantly improved, the model can still struggle with precise text placement and clarity. Second, it lists consistency challenges for recurring characters and brand elements, and composition control difficulties in layout-sensitive designs.

It also says complex prompts may take up to 2 minutes to process, which is relevant because a layout-heavy prompt with many references is exactly the kind that runs long. On Sume a request that exceeds the 30-second blocking budget returns 202 with a job envelope instead of the image body.

How do you build the wireframe?

Make it boring and exact. A flat image at the final aspect ratio, a light background, and labelled rectangles: HEADLINE, PRODUCT, LOGO, CTA. Use the labels in your prompt so the model can tie each box to an instruction. Host the file at a public HTTPS URL, because Sume rejects localhost, private-network and non-HTTPS references before submission.

Keep the final size in mind from the start. For ChatGPT Image 2.5 custom sizes, both edges must be multiples of 16, the maximum edge is 3840, the aspect ratio can be at most 3:1, and the total must fall between 655,360 and 8,294,400 pixels. A wireframe at the same ratio avoids the model inventing a new crop.

What does the request look like?

Number the references in the prompt and say what each one is for. The wireframe sets structure only; the product photo sets appearance.

{
  "model": "openai/gpt-image-2.5",
  "prompt": "Image 1 is a layout wireframe: follow its box positions exactly, but do not draw the boxes or their labels. Put the product from image 2 in the PRODUCT box. Headline in the HEADLINE box, reading exactly: \"Spring Sale\". Keep the LOGO box empty.",
  "input_references": [
    {"type": "image_url", "image_url": {"url": "https://example.com/wireframe-4x5.png"}},
    {"type": "image_url", "image_url": {"url": "https://example.com/product.png"}}
  ],
  "aspect_ratio": "4:5",
  "quality": "high"
}

What do you do when one box is still wrong?

Do not regenerate the whole ad. ChatGPT Image 2.5 on Sume takes an optional mask_url for edits, and OpenAI's guide describes the mask rules: the image and mask must match in format and size, both must be under 50 MB, the mask must contain an alpha channel, and if you supply several input images the mask applies to the first one. The guide also says masks are guidance, so the model may not follow the exact shape.

That last point matters for layout work. Treat the mask as a strong hint and check the neighbouring pixels afterwards. The edit prompt should repeat what must not change, such as the product, the background and the other boxes, before it says what to change in the masked region.

Which levers help most?

Draft at a lower quality while you tune the wireframe, then render the final at a higher one. The Sume docs list quality values auto, low, medium, high, xhigh and max, and note that auto reserves the cost of max, so pick an explicit value if you want a predictable hold.

Layout problems and the lever to try first (read 2026-10-03)
ProblemFirst leverWhere it is documented
Elements in the wrong placesWireframe as image 1, numbered in the promptSume Image API, references
Headline misplaced or garbledQuote the exact copy; fix one region with mask_urlOpenAI guide, Sume Image API
Wrong cropMatch the wireframe to the output aspect ratioSume Image API, size rules
Brand element drifts between versionsReuse the same reference every timeOpenAI guide, limitations
Final quality too lowRaise quality for the final renderSume Image API, quality

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume