FLUX Virtual Try-On v2: 4 MP inputs and a Sume reference edit

BFL's vto-v2 keeps inputs up to 4 MP as-is. Sume has no try-on endpoint in these docs; a garment swap is a reference edit with public HTTPS images.

4 min readSume
All posts

FLUX Virtual Try-On v2 (vto-v2) uses model and garment images up to 4 megapixels as-is, up from 2 in v1, with the same request format. The Sume images docs do not list a try-on endpoint; the closest path is a reference edit with input_references on an image model such as black-forest-labs/flux.2-pro.

BFL facts are from its try-on docs and release notes; Sume facts from Image generation, read 2026-10-01.

What changed in vto-v2?

BFL's release notes call vto-v2 the recommended version, focused on keeping the person's identity intact with gains in garment fidelity. Input images up to 4 MP are used as-is and the output follows the model image's resolution. The request and response format match vto-v1, so migration is a change of endpoint path (POST /v1/flux-tools/vto-v2).

What are the input limits?

Inputs above 4 MP are not rendered at full resolution: they are downscaled to about 1 MP while keeping aspect ratio. BFL suggests keeping both images near 1 MP for the best quality and latency balance, noting larger inputs raise latency.

Try-on input sizes from BFL docs, read 2026-10-01: https://docs.bfl.ml/flux_tools/flux_vto.md
VersionInput handling
vto-v2Up to 4 MP used as-is; above that, downscaled to about 1 MP
vto-v1Up to 2 MP; larger inputs downscale to about 1 MP

What does Sume accept for a garment swap?

Sume's catalog includes black-forest-labs/flux.2-pro. References go in input_references as an array of image URLs, which must be public HTTPS; localhost, private-network and non-HTTPS URLs are rejected. Whether a model takes references depends on its catalog descriptor. Put the person photo and the garment in as references and describe the swap in the prompt. This is a general edit, not BFL's try-on tool, so its output has no stated link to vto-v2 behavior.

How do I keep the person's framing?

On edit and image-to-image calls the docs recommend aspect_ratio: "auto" to match the reference; leaving the field out is not the same. For video try-on rather than stills, see virtual try-on video.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume