AI outfit builder: three garments, one model, reference order

Put a top, trousers and shoes from three separate product photos onto one model in a single gpt-image-2.5 edit on Sume, with a reference order that holds.

6 min readSume
All posts

Send the model photo as image 1 and each garment as the next images in a single POST /v1/images call to openai/gpt-image-2.5, then name each image by number and by role in the prompt: image 1 is the person, image 2 is the top, image 3 is the trousers, image 4 is the shoes. The Sume docs list up to 16 references on this model, so a full outfit fits in one request with room to spare.

ChatGPT's new Try on button covers clothing and accessories one listing at a time (read 2026-10-03, OpenAI help page). A shop that sells outfits rather than single items needs the combined view, and this is how to build it with your own product photos.

Why numbered roles beat descriptions

The model sees pixels and text, not file names. A prompt like "put the blue trousers on her" makes it guess which of three images holds the blue trousers. A prompt that says "image 3" removes the guess. Keep the order stable between calls: person, then upper body, then lower body, then footwear, then accessories. If you generate outfits in bulk, a stable order also makes your own logs readable.

Say what must not change for each item. A print on the top, the wash of the trousers and the colour of the shoes are the details a shopper will compare against the product page, so list them.

Reference order for an outfit call, read 2026-10-03
PositionRoleWhat to state in the prompt
Image 1PersonKeep face, pose and background
Image 2TopKeep print, neckline and sleeve length
Image 3TrousersKeep wash, rise and length
Image 4ShoesKeep colour and sole shape
Image 5 (optional)AccessoryKeep size relative to the body

The request

Reference URLs must be public HTTPS. Use aspect_ratio: "auto" so the output follows the person photo, or 4:5 if the destination is an Instagram portrait.

{
  "model": "openai/gpt-image-2.5",
  "prompt": "Image 1 is the model. Image 2 is the cream knit top, image 3 the black wide-leg trousers, image 4 the white sneakers. Dress the model in all three items, full body, standing, plain studio wall. Keep every item's colour, shape and details as shown.",
  "aspect_ratio": "auto",
  "quality": "high",
  "input_references": [
    {"type": "image_url", "image_url": {"url": "https://cdn.example.com/m/01.jpg"}},
    {"type": "image_url", "image_url": {"url": "https://cdn.example.com/p/top.jpg"}},
    {"type": "image_url", "image_url": {"url": "https://cdn.example.com/p/trousers.jpg"}},
    {"type": "image_url", "image_url": {"url": "https://cdn.example.com/p/sneakers.jpg"}}
  ]
}

When one call is not enough

Three or four items is usually fine. If an item gets lost, such as shoes that are cropped out of frame, say "full body, shoes visible" or run a second pass that edits only that item. The mask field mask_url is accepted on 2.5 edits as an optional public HTTPS mask, which helps for a targeted fix; check the API reference for the mask convention before you build one.

For a clip instead of a still, take the finished outfit image and continue with the seed-frame post. For a hands-off route that takes a person and one garment and returns a video, the Formats attachments post shows how many images a run can carry.

A last practical point on cost: each extra reference adds input image tokens, which the Sume docs price at $8 per million at the underlying Fal rate. Four references cost little next to the output, but if you are rendering thousands of outfits, test the price on a small batch first and read usage.cost on the response.

Checking an outfit render

An outfit has more ways to go wrong than a single garment, because the items interact. A long top can hide the waistband of the trousers, a hem can cover shoes, and layered items can merge into one. Look at each junction: neck and collar, waist, hem and ankle. If an item changed colour next to another, tighten the wording for that item and re-run at a lower quality first.

Keep a small table of results while you tune the prompt: which order of references you used, which wording, and what failed. Two or three rounds usually settle a template, and after that the same prompt works across the catalogue with only the image URLs changed.

Do not present the result as the shopper's own fit unless the person photo is theirs and they asked for it; for a generic model, say so on the page.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume