How AI virtual try-on actually works
AI virtual try-on does not overlay a picture of a garment onto a picture of you. It generates an entirely new photograph in which you are wearing it — reconstructing fabric, drape, shadow and body shape together, because in a real photograph none of those things are separable.
Almost every explanation of virtual try-on you will find describes the wrong technology. It describes warping: take a flat product image, stretch it to the shape of a body, paste it over the photo. That is how try-on worked for about a decade, and it is why for about a decade it looked like a sticker.
Modern try-on works differently, and the difference is worth understanding — because it explains both why the results are suddenly convincing and where they still go wrong.
The old way: warp and paste
The classical pipeline had three stages. Segment the person in the photo to work out which pixels are body and which are the clothes they are already wearing. Warp the flat garment image onto the estimated body shape using a geometric transform. Composite the warped garment over the segmented region and blend the edges.
It works, in the narrow sense that something appears on the body. It fails in every way that matters. A geometric warp has no idea that a heavy wool coat hangs differently from a silk slip — it stretches both the same. It cannot generate a fold that was not in the source image, so a garment photographed flat stays flat. It has no lighting model, so the garment carries the studio lighting of the product shoot into a photo taken in your bedroom. And because it composites, the boundary between the garment and the arm is a hard edge that no amount of feathering fixes.
The tell of a warped try-on is always the same: the clothes are lit differently from the person wearing them.
The new way: generate the whole photograph
Generative try-on abandons compositing entirely. Instead of moving pixels from a product image onto a person image, a diffusion model is asked to produce a new image from scratch, conditioned on two things at once: this person, and this garment.
Diffusion models generate by denoising. The model starts from pure noise and, over a sequence of steps, repeatedly predicts what the noise is and removes a little of it, until an image emerges. What makes the output controllable is what the model is allowed to look at while it does this. In try-on, it looks at a representation of your body and a representation of the garment simultaneously, and every denoising step is nudged toward an image consistent with both.
The consequence is the whole point: because the image is generated rather than assembled, the fabric, the folds, the shadow the collar casts on the neck and the way the hem breaks over a hip are all produced together, by one process, under one lighting model. Nothing has to be blended, because nothing was ever separate.
Keeping your face, and the rest of you
A generative model producing a photograph of a person will, left alone, produce a photograph of a person. Making sure it produces one of you is the hardest part of the problem, and it is where most try-on demos quietly fail — the results look great and the face is subtly someone else's.
The approach that works is to separate the identity from the generation. Before any try-on happens, the system builds a persistent representation of you from your photo — Utopia calls this a fit model — which encodes face, body proportions, skin tone and hair as a reusable conditioning signal. Every subsequent generation is anchored to that same representation.
This is why building the fit model takes longer than a try-on does (roughly 40 to 70 seconds, once) and why every try-on afterwards is faster (roughly 30 to 60 seconds). The expensive part happens a single time. It is also why the twentieth try-on looks like the same person as the first, which a per-image approach cannot guarantee.
What the model has to get right
A convincing result is not one thing. It is several independent problems that all have to be solved in the same image:
- Garment identity. The specific piece, not a generic version of it. The print has to be the actual print, at the correct scale, in the correct place — a repeating pattern that drifts across a fold is the most common giveaway.
- Material behaviour. Wool, denim, silk and jersey fall differently under gravity. The model has to infer the material from a product photograph and then simulate its consequences.
- Body geometry. The garment has to be occluded correctly by arms, hair and anything the person is holding, and it has to change shape where the body underneath it changes shape.
- Lighting coherence. One light source for the whole image. If the garment is lit from the left and the face from the right, the eye catches it instantly even if it cannot say why.
- Identity preservation. Same face, same proportions, same skin tone, every time.
Where it still goes wrong
It would be easy to write this section as a list of things that are nearly solved. They are not, and knowing the failure modes makes the output far more useful:
- Fine text and small logos. A generated image reconstructs a logo from a low-resolution understanding of it. Large graphics survive; five-point type on a chest pocket does not.
- Very complex hardware. Buckles, chain straps and layered fastenings are where the reconstruction is most likely to invent a detail that is not on the real garment.
- Extreme poses. A crossed-arms, three-quarter-turned, one-leg-raised photo asks the model to infer geometry it cannot see. A relaxed, front-facing photo is dramatically more reliable — the reason how you take the photo matters so much.
- Size, as opposed to fit. This is the important one. A try-on shows you how a garment falls on your proportions. It is not a measurement, and it cannot tell you that you are between a medium and a large.
Why it took this long
Virtual try-on has been an active research area since well before generative image models were any good, and the reason it stayed unconvincing is that it was framed as a graphics problem — estimate the geometry, warp the texture, composite the layers. Framed that way, every step introduces an error the next step has to hide.
Reframing it as an image-generation problem removed the compositing step entirely, and with it most of the artefacts. The remaining difficulty moved from geometry to control: how do you generate freely enough to produce realistic cloth, while constraining tightly enough that the person is unmistakably the right person wearing the exact right garment? That is where the engineering effort now goes.
Trying it on something real
The description above is only worth so much. The catalogue in Utopia runs to more than 400,000 pieces from close to 300 brands, and the categories where generated try-on is most obviously better than a product photo are the ones where drape does the work — a puffer, where loft changes your whole outline, denim, where rise decides everything, or an unornamented COS silhouette that is judged entirely on how it sits.
Add one photo and see the difference between a warp and a generated image on your own body.
Get the appFrequently asked
How does AI virtual try-on work?
+
AI virtual try-on generates a new photograph in which you are wearing a garment, rather than pasting the garment over an existing photo. A diffusion model denoises an image while being conditioned on both a representation of your body and a representation of the garment, so fabric, folds, shadow and lighting are produced together instead of composited.
Is AI try-on accurate?
+
It is accurate about fit and drape — how a garment falls on your proportions — and it is not a measurement. Generated try-on reliably reproduces silhouette, fabric behaviour and colour; it is less reliable on fine text, small logos and complex hardware, and it cannot tell you which size to order.
How long does an AI try-on take?
+
In Utopia, building your fit model from a photo takes roughly 40 to 70 seconds and happens once. Each try-on after that takes roughly 30 to 60 seconds, because the expensive identity work has already been done.
What is the difference between virtual try-on and an AR filter?
+
An AR filter tracks a live camera feed and overlays a flat graphic in real time, so it slides as you move and carries no fabric behaviour. Generative virtual try-on produces a single still photograph with real drape, shadow and lighting, which you can save and compare.
See it on you, not on a model.
Utopia puts any of 400,000+ pieces on your own photo in about a minute. Free on the App Store.
Get Utopia