Try-on with front and back garment photos: three references
When a garment's back has a print the front lacks, send shopper, front and back photos as three references. Prompt wording, a Python call and what to check.

ChatGPT Try On takes a selfie or full-body photo and, per TechCrunch (read 2026-10-04), can also work from a screenshot of an item. A screenshot is one view. A tee with a large back print or a jacket with a hood detail needs two.
On Sume's Image API the GPT Image 2.5 model accepts up to 16 references, so shopper, front and back can travel together as three.
Name each reference
The order of the input_references array is the only label the model gets, so repeat it in the prompt: image 1 is the shopper, image 2 the garment front, image 3 the garment back. Then say which side the output shows. A single image cannot show front and back at once, so run two calls from the same references: one facing the camera, one turned away.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
resp = requests.post("https://api.sume.com/v1/images", headers=H, timeout=90, json={
"model": "openai/gpt-image-2.5",
"prompt": "Image 1 is the shopper. Image 2 is the front of a black hoodie. Image 3 is the back of the same hoodie with a white print. Show the shopper from behind wearing this hoodie, with the print from image 3. Keep body, pose and background from image 1.",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/shopper.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/hoodie-front.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/hoodie-back.jpg"}}
],
"aspect_ratio": "auto",
})
resp.raise_for_status()
if resp.status_code == 202:
raise SystemExit("job envelope: poll data.status_url")
body = resp.json()
print([d["url"] for d in body["data"]], body.get("usage"))Checks that catch the usual misses
- Print placement: compare height from the collar with the product photo.
- Text on the print is spelled as in image 3, letter for letter.
- Sleeve length and hem match the product's size chart for that shopper.
- The turned-away shot shows no face from image 1 that contradicts the pose.
Cost
Two calls for two views. Each response includes usage.cost, and more references add input tokens, so check that field on a sample before rolling it into a catalog. The docs give $0.09366 for a 1024 by 1024 xhigh output at fal's token rates, before input tokens and Sume's pricing, which is a ceiling worth knowing when you pick a quality tier.
Sources
Related posts
More in Use cases
- Tumbler or water bottle photos with AI: composite the logo in code
Brand logos on bottles drift in generated images. Make the scene with Sume's image API, then paste your transparent logo with Pillow so the mark is exact.
- Turn a drawing into an image with GPT Image 2.5
Wikipedia says GPT Image 2.5 added Sketch, which turns drawings into images. Sume exposes GPT Image 2.5 with image references; here is how to send a drawing.
- Turn a Short into podcast audio with audio detach for $0.01
Audio detach extracts the sound of a hosted video into a wav or mp3 for $0.01 per job. Use it to turn a clip of up to 3 minutes into a podcast cut.
- Tutor intro video: approve the first frame, then pick the tier
A tutor can approve an avatar intro's first frame before paying for the full render, then choose standard, plus or max at generate-video without a new preview.
Written by Sume