AI try-on for sunglasses, bags and scarves from a photo

ChatGPT Try On covers accessories as well as clothes. The same edit on Sume: several photos of the item as references, and what to check on scale and shape.

5 min readSume
All posts

Use an image edit with two or three photos of the item as input_references next to the person photo: front, side, and one worn if you have it. On Sume that is openai/gpt-image-2.5 on POST /v1/images, which takes up to 16 references. ChatGPT Try On is reported to cover clothes and accessories, but a generated accessory is a likeness of the product, and sunglasses, bags and scarves fail in specific ways you can check for.

Below: what the launch coverage says, the call, and a short list of checks for scale, shape and logo.

Does ChatGPT Try On handle accessories?

TechCrunch says a user can upload a selfie or full-body photo to see how an article of clothing or an accessory might look on them. The same report ties it to ChatGPT Images 2.5. Retail Technology Innovation Hub adds that OpenAI notes images may not reproduce the person or product exactly. We could not read OpenAI's own page, so both are reported statements. For a brand, the caveat is the useful part: an accessory image is a hint, so it needs your real photography behind it.

What does the accessory call look like?

One request, the item shown from more than one side, and a prompt that lists the details that must not change. The reference limit and the public-HTTPS rule come from the Image API docs. Sume does not label the references, so say in the prompt which is which, and check the result on your own pair before you run a batch.

Photograph the item on a plain background in even light. A bag shot at an angle with a cluttered shelf behind it gives the model nothing clean to copy.

curl -sS -X POST "https://api.sume.com/v1/images" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sunglasses-p1-v1" \
  -d '{
    "model": "openai/gpt-image-2.5",
    "prompt": "The first image is the person. The second and third images show the same sunglasses from the front and the side. Put the sunglasses on the person. Keep face, hair and background unchanged. Keep the frame shape, lens tint and arm width exactly as shown.",
    "input_references": [
      {"type": "image_url", "image_url": {"url": "https://cdn.example.com/person.jpg"}},
      {"type": "image_url", "image_url": {"url": "https://cdn.example.com/sunglasses-front.png"}},
      {"type": "image_url", "image_url": {"url": "https://cdn.example.com/sunglasses-side.png"}}
    ],
    "aspect_ratio": "auto"
  }'

What goes wrong with accessories?

Clothes drape; accessories have a fixed shape and a real size. That makes the errors easier to spot and harder to fix by prompting. Look at the generated image beside the packshot before you publish.

Accessory try-on checks, read 2026-10-03
ItemWhat to checkWhy it fails
SunglassesFrame shape, lens tint, arm width, how they sit on the noseA model can reshape a frame to fit the face
HandbagSize against the body, strap length, hardware and logoA generated bag has no known dimensions
ScarfPattern placement and scale, how it is knottedPrint repeats are easy to rescale
HatBrim width and crown heightA flat brim can turn into a curved one
Watch or braceletFace size against the wristSmall items shrink or grow without notice

How many views of the item does it need?

At least two. A single front view tells the model nothing about depth: how far the arms of a pair of glasses reach, how thick a bag is, how a scarf falls. The ChatGPT Image 2.5 models on Sume accept up to 16 references, so a person photo plus several views of the item is well inside the limit. More views do not guarantee a better result, but they remove guesses you can remove.

Keep the item's own photos free of people and props. If one reference shows the item on a model and another shows it flat, the prompt has to say which one is the truth for the shape and which one is only for the colour; otherwise the model will blend them.

What should the listing say?

Say what the image is: a visual preview made with AI from product photos. Do not use it to state a size, a drop length or a strap length; give those from the real item. If your platform has a field for the real product image, put the real image first and the preview after it, because the shopper's trust lives in the first picture they see.

The same applies to the consumer side. OpenAI's own reported caveat for ChatGPT Try On is that images may not reproduce the product exactly, so a seller that publishes its own preview should be at least as careful.

Is a video a better fit?

For a bag or a scarf, motion is the point: a swing, a flick, a turn. Sume's catalog has two apparel Formats, but they are written for garments, so for accessories the safer route is the one above: make and approve a still, then animate it with seedance-2.5 as a first frame, which the video docs list at 4 to 30 seconds. For which accessories a platform's own avatar tool will refuse, see the TikTok avatar exclusions.

  • Use a person photo you have the right to edit.
  • Give the model two or three views of the item, not one.
  • State the details that must not change in the prompt.
  • Compare the result with your packshot at full size.
  • Do not publish a size claim from a generated image.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume