What is virtual try-on? AR try-on and AI try-on explained
Virtual try-on shows a product on a person who isn't wearing it: live on a camera feed with AR, or in a new AI-generated image or video made from photos.

Virtual try-on is technology that shows a product on a person who isn't physically wearing it. AR try-on overlays the item live on a camera feed, so shoppers see glasses or makeup on their own face as they move. AI try-on generates a new image or video of a person wearing a garment, from a photo of the person and a photo of the garment.
The two kinds are explained below in plain terms. Where Sume fits comes from its Format catalog, Image API, and Video generation docs, read on 2026-09-28. Sume makes the second kind.
How does AR virtual try-on work?
Augmented reality try-on runs live on the shopper's own camera. The app follows a face, a hand, or a foot and draws the product onto it as the person moves: frames on the nose, lipstick on the lips, a watch on the wrist. The shopper sees themselves, and nothing needs to be saved as media. That is why it is associated with items worn on the face or hands, such as glasses, makeup, and jewelry.
How does AI virtual try-on work?
Generative try-on works from photos, not a camera. A model takes a picture of a person and a picture of the garment and generates a new picture, or a clip, of that person wearing it. The output is a file you can publish anywhere: a product page, an ad, a lookbook. It suits clothing, because the result shows the garment on a whole body in one frame.
Some AI try-on tools return a still image and others return a clip, so check which one a tool makes before you plan around it. For one vendor's image API, see Kling virtual try-on API.
| Kind | Output | Example |
|---|---|---|
| AR try-on | Live overlay, nothing saved | Camera apps for glasses and makeup |
| AI try-on still | An image | Sume POST /v1/images with the person and garment as references |
| AI try-on video | A video | Sume Formats sume-virtual-try-on and sume-virtual-fitting |
How do I make an AI try-on with Sume?
Sume makes the AI kind only, from photos at public HTTPS URLs, in two ways. For a video, call a catalog try-on Format, sume-virtual-try-on or sume-virtual-fitting; both descriptions end “Not for: static campaign deliverables”, so both make video. For a still, send the person and the garment as references to POST /v1/images, then, if you want motion, copy the still you approve to your own public HTTPS URL (image result URLs are signed) and animate it with POST /v1/videos.
Both routes, field by field, are in Virtual try-on video API. For a store, Virtual try-on for Shopify covers where the finished video goes.
What can't AI virtual try-on do?
It can't see the shopper. Every viewer gets the same generated person, so it doesn't answer “how does this look on me” the way a live camera does. It also doesn't measure anything: “accurate garment fit” is the fitting Format's stated aim, not a sizing tool, so check each output against the real garment. Use photos of people who agreed to appear.
Sume's try-on routes work from supplied photos, not a live camera feed, so they make media rather than a live try-on feature.
Sources
Related posts
More in Use cases
- What is YouTube automation? How it works with AI
YouTube automation is running a channel, often faceless, whose production goes to freelancers or AI tools. What AI does, and what stays yours.
- YouTube AI content policy: what's allowed, what to disclose
YouTube allows AI videos but requires you to disclose realistic AI-generated or altered content, and won't monetize mass-produced, templated AI videos.
- YouTube end screen template: a 5–20 second 16:9 outro
A YouTube end screen template is a 16:9 outro of 5–20 seconds with room for up to four elements. YouTube's rules, and how to make the outro with AI.
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
Written by Sume