ChatGPT Try On from a screenshot: the same edit through an API

ChatGPT Try On starts from a selfie plus a product screenshot. Do the same edit with openai/gpt-image-2.5 on Sume: two references, one prompt, one Python call.

5 min readSume
All posts

You can run the same kind of edit by API: send the person photo and the garment screenshot as two input_references to openai/gpt-image-2.5 on POST /v1/images, and say in the prompt which image is the person and which is the garment. ChatGPT Try On is a consumer feature, and none of the coverage we read describes an API for it, so this is the closest documented route on Sume, not a clone of OpenAI's product.

OpenAI says Try On runs on ChatGPT Images 2.5, which Sume lists as openai/gpt-image-2.5 (Flare) and openai/gpt-image-2.5-sunburst.

What does ChatGPT Try On start from?

The Neuron's October 1 digest, which we cite only as a report, describes it as: upload a selfie, a full-body photo or a product screenshot, and Images 2.5 generates the outfit on you and saves the result to a Library. TechCrunch adds that you can upload an image of an item, such as a web screenshot, and ask ChatGPT to try it on. So the inputs are two pictures and an instruction, which is exactly the shape of an image edit.

What does Sume's image call accept?

Per the Image API docs, the ChatGPT Image 2.5 models support text-to-image and up to 16 image references. References must be public HTTPS URLs; localhost, private-network and non-HTTPS URLs are rejected before submission. On edits the docs advise aspect_ratio: "auto" to match the reference, and they add that omitting the field is not the same as auto.

Try-on inputs mapped to the Sume image call, read 2026-10-03
You haveSume fieldRule from the docs
Selfie or model photoinput_references[0]Public HTTPS URL
Garment screenshot or packshotinput_references[1]Public HTTPS URL; up to 16 references in total
Which image is whichpromptWritten by you; Sume does not label references
Keep the framingaspect_ratio: autoPrefer on edits
Output size and costquality, image_sizeDefault quality is high; read the endpoint's pricing first

What does the call look like?

The call blocks for up to 30 seconds and returns 200 with data[].url; a slow one degrades to 202 with a job envelope, so check the status code rather than the body shape. This script handles both and sends an Idempotency-Key so a retried POST is not a second charge.

Host both pictures first. A ChatGPT screenshot on your phone is not a URL until you upload it somewhere public.

import os
import requests

key = os.environ["SUME_API_KEY"]
ref = lambda u: {"type": "image_url", "image_url": {"url": u}}
body = {
    "model": "openai/gpt-image-2.5",
    "prompt": (
        "The first image is the person, the second is the garment. "
        "Dress the person in that garment. Keep face, pose, background "
        "and lighting unchanged. Keep the garment's colour, print and "
        "sleeve length exactly as shown."
    ),
    "input_references": [
        ref("https://cdn.example.com/person.jpg"),
        ref("https://cdn.example.com/shirt.png"),
    ],
    "aspect_ratio": "auto",
    "output_format": "png",
}
headers = {"Authorization": f"Bearer {key}", "Idempotency-Key": "tryon-p1-shirt1-v1"}
r = requests.post("https://api.sume.com/v1/images", headers=headers, json=body, timeout=60)
if r.status_code == 200:
    print(r.json()["data"][0]["url"])
elif r.status_code == 202:
    print("still running:", r.json()["data"]["status_url"])
else:
    raise SystemExit(r.text)

What should you check before you trust the result?

A prompt that names the images by order is a habit, not a documented guarantee, so test it once on your own pair and look at the result before you batch. Then check the same things a shopper will: the print's placement, the colour against your packshot, the sleeve length, and the hands. Each failed generation is not billed, per the docs, while a completed one is billed in full, so a bad result still costs money; fix the prompt, not the retry count.

  • Use a photo of a person you have the right to edit; the call has no consent check.
  • Use a flat, evenly lit garment image so the colour is readable.
  • Re-run with one change at a time so you can tell what moved the result.
  • Keep the original garment photo beside the generated one in review.

How is this different from what ChatGPT does for a shopper?

ChatGPT keeps your reference photo for next time and shows a Try On button next to listings, reported by Retail Technology Innovation Hub. An API call has none of that: no saved profile, no shopping graph, no button. You hold the person photo, you hold the garment image, and you write the instruction every time. What you get back is a Sume-hosted URL per image in data[].url rather than a Library entry.

That is a fair trade for a brand. You decide which people appear, which garments, how the result is framed, and where it is published. If you want the same garment on a creator in motion instead of a still, Sume's catalog has sume-virtual-try-on, which is a video Format; the still route above is the right one when the deliverable is a picture for a product page.

One cost note, from the docs only. The Image API says a completed generation is billed in full and a failed one is not, and that the endpoint's pricing lines are what your wallet is charged. Read GET /v1/images/models/openai/gpt-image-2.5/endpoints before you loop over a catalog, because a quality setting of auto reserves the max price.

Sources

Related posts

More in Models

All Models posts

Written by Sume