Try-on with front and back garment photos: three references

When a garment's back has a print the front lacks, send shopper, front and back photos as three references. Prompt wording, a Python call and what to check.

4 min readSume
All posts

ChatGPT Try On takes a selfie or full-body photo and, per TechCrunch (read 2026-10-04), can also work from a screenshot of an item. A screenshot is one view. A tee with a large back print or a jacket with a hood detail needs two.

On Sume's Image API the GPT Image 2.5 model accepts up to 16 references, so shopper, front and back can travel together as three.

Name each reference

The order of the input_references array is the only label the model gets, so repeat it in the prompt: image 1 is the shopper, image 2 the garment front, image 3 the garment back. Then say which side the output shows. A single image cannot show front and back at once, so run two calls from the same references: one facing the camera, one turned away.

import os, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
resp = requests.post("https://api.sume.com/v1/images", headers=H, timeout=90, json={
    "model": "openai/gpt-image-2.5",
    "prompt": "Image 1 is the shopper. Image 2 is the front of a black hoodie. Image 3 is the back of the same hoodie with a white print. Show the shopper from behind wearing this hoodie, with the print from image 3. Keep body, pose and background from image 1.",
    "input_references": [
        {"type": "image_url", "image_url": {"url": "https://example.com/shopper.jpg"}},
        {"type": "image_url", "image_url": {"url": "https://example.com/hoodie-front.jpg"}},
        {"type": "image_url", "image_url": {"url": "https://example.com/hoodie-back.jpg"}}
    ],
    "aspect_ratio": "auto",
})
resp.raise_for_status()
if resp.status_code == 202:
    raise SystemExit("job envelope: poll data.status_url")
body = resp.json()
print([d["url"] for d in body["data"]], body.get("usage"))

Checks that catch the usual misses

  • Print placement: compare height from the collar with the product photo.
  • Text on the print is spelled as in image 3, letter for letter.
  • Sleeve length and hem match the product's size chart for that shopper.
  • The turned-away shot shows no face from image 1 that contradicts the pose.

Cost

Two calls for two views. Each response includes usage.cost, and more references add input tokens, so check that field on a sample before rolling it into a catalog. The docs give $0.09366 for a 1024 by 1024 xhigh output at fal's token rates, before input tokens and Sume's pricing, which is a ceiling worth knowing when you pick a quality tier.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume