Jacket try-on over the shopper's own outfit: prompt wording that holds
A jacket goes over clothes, not instead of them. Prompt wording and a two-reference Sume image call that keep the shopper's shirt visible, plus layering checks.

Most try-on demos swap one garment for another. A coat or jacket is different: the shopper wears something under it. ChatGPT Try On, launched 1 October 2026 for clothing and accessories (TechCrunch, read 2026-10-04), starts from a full-body photo, so the outfit in the photo is the base the new item must layer on.
In the Sume Image API the same thing is a two-reference edit, and the prompt does the layering work.
Say what stays
Default edits tend to replace the nearest garment. Name what must remain: the shirt, its colour, the neckline showing at the collar, and the trousers. Name what changes: the jacket only, worn open or closed as you want it.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
resp = requests.post("https://api.sume.com/v1/images", headers=H, timeout=90, json={
"model": "openai/gpt-image-2.5",
"prompt": "Image 1 is the shopper in their own outfit. Image 2 is a tan field jacket. Put the jacket on over the shopper's existing shirt, worn open so the shirt shows at the chest. Keep the shirt, trousers, pose, face and background from image 1. Keep the jacket's pockets and buttons as in image 2.",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/shopper-full.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/jacket.jpg"}}
],
"aspect_ratio": "auto",
})
resp.raise_for_status()
if resp.status_code == 202:
raise SystemExit("job envelope: poll data.status_url")
body = resp.json()
print([d["url"] for d in body["data"]], body.get("usage"))Layering checks
- The shirt colour at the collar and hem is the same as in the base photo.
- The jacket's shoulder seam sits on the shoulder, not above it.
- Cuffs show the shopper's real wrist, not a merged sleeve.
- Open versus closed matches what you asked for.
When one reference is not enough
If the jacket has a lining you want to show, add a third reference with the lining visible and name it in the prompt. Reference URLs must be public HTTPS, and Sume's docs allow up to 16 on GPT Image 2.5. Keep aspect_ratio at auto so the output matches the shopper's photo, and read usage.cost on every call to know what a preview costs you.
Before you run a catalog
Run one SKU end to end first. POST /v1/images returns the images directly when it finishes within 30 seconds; past that it returns a 202 envelope with status_url and result_url, which the jobs and results guide explains. Handle that branch before you loop over a catalog, and write each result's URL and usage.cost to a file keyed by SKU, so a failed run restarts where it stopped and nothing is paid for twice.
Sources
Related posts
More in Use cases
- Keep one character consistent in a 30-second Seedance 2.5 clip
Character consistency in a single 30 s Seedance 2.5 pass: send reference_image_urls on Sume, test at 480p, and judge stills. No guarantee, a cheap test plan.
- Kettle or air fryer product photos: cord, spout and counter scene
Kitchen appliances fail on cords, handles and dials. Generate counter scenes with Sume's image API from one reference and check the parts that go wrong first.
- Kling 4.0 keyframes at 0, 6, 12, 20, 28 s: a 30-second ad on Sume
Kling's guide suggests keyframes near 0, 6, 12, 20 and 28 seconds for a product ad. How to build the same plan on Sume with first-and-last-frame clips.
- Kling 4.0 lip-sync reaches nine languages: a one-clip test each
Kling 4.0 adds Portuguese, German, French and Hindi lip-sync to five existing languages. A short per-language test plan to run before you promise localization.
Written by Sume