AI try-on for sunglasses, bags and scarves from a photo
ChatGPT Try On covers accessories as well as clothes. The same edit on Sume: several photos of the item as references, and what to check on scale and shape.

Use an image edit with two or three photos of the item as input_references next to the person photo: front, side, and one worn if you have it. On Sume that is openai/gpt-image-2.5 on POST /v1/images, which takes up to 16 references. ChatGPT Try On is reported to cover clothes and accessories, but a generated accessory is a likeness of the product, and sunglasses, bags and scarves fail in specific ways you can check for.
Below: what the launch coverage says, the call, and a short list of checks for scale, shape and logo.
Does ChatGPT Try On handle accessories?
TechCrunch says a user can upload a selfie or full-body photo to see how an article of clothing or an accessory might look on them. The same report ties it to ChatGPT Images 2.5. Retail Technology Innovation Hub adds that OpenAI notes images may not reproduce the person or product exactly. We could not read OpenAI's own page, so both are reported statements. For a brand, the caveat is the useful part: an accessory image is a hint, so it needs your real photography behind it.
What does the accessory call look like?
One request, the item shown from more than one side, and a prompt that lists the details that must not change. The reference limit and the public-HTTPS rule come from the Image API docs. Sume does not label the references, so say in the prompt which is which, and check the result on your own pair before you run a batch.
Photograph the item on a plain background in even light. A bag shot at an angle with a cluttered shelf behind it gives the model nothing clean to copy.
curl -sS -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sunglasses-p1-v1" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "The first image is the person. The second and third images show the same sunglasses from the front and the side. Put the sunglasses on the person. Keep face, hair and background unchanged. Keep the frame shape, lens tint and arm width exactly as shown.",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://cdn.example.com/person.jpg"}},
{"type": "image_url", "image_url": {"url": "https://cdn.example.com/sunglasses-front.png"}},
{"type": "image_url", "image_url": {"url": "https://cdn.example.com/sunglasses-side.png"}}
],
"aspect_ratio": "auto"
}'What goes wrong with accessories?
Clothes drape; accessories have a fixed shape and a real size. That makes the errors easier to spot and harder to fix by prompting. Look at the generated image beside the packshot before you publish.
| Item | What to check | Why it fails |
|---|---|---|
| Sunglasses | Frame shape, lens tint, arm width, how they sit on the nose | A model can reshape a frame to fit the face |
| Handbag | Size against the body, strap length, hardware and logo | A generated bag has no known dimensions |
| Scarf | Pattern placement and scale, how it is knotted | Print repeats are easy to rescale |
| Hat | Brim width and crown height | A flat brim can turn into a curved one |
| Watch or bracelet | Face size against the wrist | Small items shrink or grow without notice |
How many views of the item does it need?
At least two. A single front view tells the model nothing about depth: how far the arms of a pair of glasses reach, how thick a bag is, how a scarf falls. The ChatGPT Image 2.5 models on Sume accept up to 16 references, so a person photo plus several views of the item is well inside the limit. More views do not guarantee a better result, but they remove guesses you can remove.
Keep the item's own photos free of people and props. If one reference shows the item on a model and another shows it flat, the prompt has to say which one is the truth for the shape and which one is only for the colour; otherwise the model will blend them.
What should the listing say?
Say what the image is: a visual preview made with AI from product photos. Do not use it to state a size, a drop length or a strap length; give those from the real item. If your platform has a field for the real product image, put the real image first and the preview after it, because the shopper's trust lives in the first picture they see.
The same applies to the consumer side. OpenAI's own reported caveat for ChatGPT Try On is that images may not reproduce the product exactly, so a seller that publishes its own preview should be at least as careful.
Is a video a better fit?
For a bag or a scarf, motion is the point: a swing, a flick, a turn. Sume's catalog has two apparel Formats, but they are written for garments, so for accessories the safer route is the one above: make and approve a still, then animate it with seedance-2.5 as a first frame, which the video docs list at 4 to 30 seconds. For which accessories a platform's own avatar tool will refuse, see the TikTok avatar exclusions.
- Use a person photo you have the right to edit.
- Give the model two or three views of the item, not one.
- State the details that must not change in the prompt.
- Compare the result with your packshot at full size.
- Do not publish a size claim from a generated image.
Sources
Related posts
More in Use cases
- AI virtual try-on looks fake? A QC checklist for print and hands
AI try-on can drift from the real garment. Check print, colour, sleeves and hands, and pull exact frames from a clip with Sume's video frames endpoint.
- AI voice reading news articles on YouTube Shorts: allowed?
YouTube lists readings of websites or news feeds as unoriginal material. How to turn a news story into an original Short with your own analysis.
- AI workout music for fitness videos: tempo and intervals in a prompt
Make AI workout music by writing tempo as a number and the interval plan as timestamps. Sume has no BPM field, so here is how to prompt, check and trim a track.
- Animate a locally generated image with Sume video: first frame
Made a still with a local open-weights model? Host it at a public HTTPS URL, send it as first_frame to POST /v1/videos, and poll. Plus the licence check first.
Written by Sume