One garment on different body types with AI: use real people
Shoppers want clothes on bodies like theirs. Google's 2023 try-on showed real models XXS to 4XL. Build a range from consented photos and one Sume image call.

Photograph a small set of real people who agreed to appear, in sizes that reflect your customers, and run the same garment through the same image edit for each one. On Sume that is one POST /v1/images call per person with openai/gpt-image-2.5: the person photo and the garment photo as input_references. You get the range of bodies shoppers ask for without asking an image model to invent people, and every face in your listing belongs to someone you can name.
The reason to do this now is that try-on has moved into the shopping assistant. ChatGPT Try On launched on October 1 2026 and is reported to make no size or fit promise, so the sellers' own pictures carry more of the weight.
Why does body range matter in the first place?
Google's launch post for its earlier try-on cites a survey it gives as 42 percent of online shoppers not feeling represented by images of models, and 59 percent being dissatisfied with something that looked different on them than expected. It said its model shows a garment on real models in sizes XXS to 4XL, with a range of skin tones, body shapes, ethnicities and hair types (Google, page dated June 14 2023, read 2026-10-03). The statistics are Google's, from 2023, and we cannot check them; the direction they point is why sellers show more than one model.
What is the safe way to build the set?
Cast people, get written permission for the use, and keep the originals. A generated image of an invented person is a different thing: it can look like a real person by accident, and nothing in Sume's image call checks. The consent work is yours, and the digital replicas checklist is a starting point for the law in California, not legal advice.
| Step | What you do | Sume part |
|---|---|---|
| Cast | Real people across your size range, with written permission | None; this is yours |
| Photograph | Same pose, plain light, full length | None |
| Host | Public HTTPS URL per photo | Image references must be public HTTPS |
| Edit | One garment image plus one person image per call | openai/gpt-image-2.5, up to 16 references |
| Review | Compare each result with the real garment and the person | None; a human step |
How do you run it for each person?
Loop over your people with the same garment and the same prompt, and give each call an idempotency key built from the garment and the person. A sync call returns a Sume-hosted URL in data[0].url, per the Image API docs; a slow one returns a 202 job instead, which the script prints.
Use identical framing for the originals so the results line up on a grid.
import os
import requests
key = os.environ["SUME_API_KEY"]
headers = {"Authorization": f"Bearer {key}"}
garment = "https://cdn.example.com/dress.png"
people = {
"p1": "https://cdn.example.com/people/p1.jpg",
"p2": "https://cdn.example.com/people/p2.jpg",
}
ref = lambda u: {"type": "image_url", "image_url": {"url": u}}
for name, url in people.items():
body = {
"model": "openai/gpt-image-2.5",
"prompt": "First image: the person. Second image: the dress. Put the dress on the person. Keep face, body shape, pose and background. Keep the dress's colour and print.",
"input_references": [ref(url), ref(garment)],
"aspect_ratio": "auto",
}
r = requests.post("https://api.sume.com/v1/images", headers={**headers, "Idempotency-Key": f"dress1-{name}-v1"}, json=body, timeout=60)
print(name, r.status_code, r.json()["data"][0]["url"] if r.status_code == 200 else r.text[:200])How many people are enough?
Enough that a shopper can find someone near their own height and build, and few enough that you can review every image by hand. For a first set that can be as few as four or five, photographed once, then reused for every new garment. Each new garment is one call per person, so the work grows with your range, not with a studio booking.
Keep a record per person: who they are, what you agreed, which uses it covers, and the date. Re-ask when the use changes, for example when a still becomes a video.
What should you not claim?
Do not say a generated picture shows how a size will fit. The image model is told a garment and a body; it does not know your pattern or your fabric. Say the picture is a visual preview, give the model's height and the size they wear in the caption, and keep your size chart next to it. If a platform asks for an AI label on the image, apply it.
A Format is the other option: sume-virtual-fitting makes a 9:16 clip for silhouette and drape, per the Formats catalog, and the same rules apply to it.
- Get written permission from each person, for each use.
- State the model's height and worn size in the caption.
- Keep the real garment photo next to the generated image.
- Do not generate a person who does not exist and present them as a customer.
- Label AI images where a platform requires it.
Sources
Related posts
More in Use cases
- Snap DPA for Sponsored Snaps: catalog images to fix with Sume
Snap's Dynamic Product Ads now run in Chat for all global advertisers. The creative comes from your catalog, so the work is catalog images Sume can fix.
- Swap multiple characters in a video with AI: 2 to 4 people
Swap several people in one clip with H3 Max Recast on Sume: one photo per person, left to right by default, up to four, and a price that ignores headcount.
- Template bulk edits on YouTube Shorts: what to vary per row
YouTube's Oct 1, 2026 originality update names template-based bulk changes as not original. How to make each row of a Sume bulk run differ in substance.
- Where GMV Max creative videos come from, and where AI clips go
TikTok Product GMV Max only runs videos from connected accounts, Spark Ads posts or ACA, with a product anchor link. Where a Sume-made MP4 has to go first.
Written by Sume