AI virtual try-on cost per image: GPT Image 2.5 on Sume
What one try-on still costs on Sume: the documented token rates, why quality auto reserves the max price, and a script that reads the live endpoint pricing.

A try-on still on Sume is one openai/gpt-image-2.5 call, and its price is what the endpoint's pricing lines say at the quality you choose, not a flat per-try-on fee. Sume's docs give the underlying rates: $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens, and at 1024 by 1024 an xhigh output is $0.09366 and max is $0.21072 before input tokens and Sume pricing. A try-on call sends at least two input images, so input tokens count too.
ChatGPT Try On, launched October 1 2026 as a shopping feature, is not priced per image in anything we read. This post is about the API route.
Where do those numbers come from?
They are in the Image API docs, which say Flare and Sunburst use the same Fal token rates and link to the Fal model page. The docs add that output estimates use OpenAI's ChatGPT Image 2.5 size and quality calculator, that input token counts are estimates, and that Fal rounds the total up to $0.0001. We did not open the Fal page, so the rates here are exactly as Sume's docs state them.
What does a try-on call add to the output price?
Three things. First, input images: a try-on sends a person photo and a garment photo, and input image tokens are billed at the input rate. Second, Sume's margin: the docs say an endpoint's pricing lines are the amount charged to your wallet, with the margin already applied. Third, reservation: auto quality reserves max, and auto size and named presets without a verified pixel mapping reserve the output token upper bound. Omitted quality defaults to high.
| Factor | Documented rule | What to do |
|---|---|---|
| Output rate | $30 per million output image tokens | Pick quality deliberately |
| Input images | $8 per million input image tokens; counts are estimates | Send only the references the edit needs |
| 1024x1024 xhigh output | $0.09366 before input and Sume pricing | Use as a floor for a square test |
| 1024x1024 max output | $0.21072 before input and Sume pricing | Avoid unless a print needs it |
| quality: auto | Reserves the max price | Set quality explicitly |
| Failed generation | Not billed; a completed one is billed in full | Check results before looping |
How do you read the live price before you batch?
Ask the endpoint. GET /v1/images/models/{model_id}/endpoints returns the billable lines for the model, per the docs example. The script below prints them; run it before you queue a catalog, and again if the price might have changed. The response example in the docs has no data wrapper, so the script accepts either shape.
Then run one real try-on call and read usage.cost from the response, which the docs define as the billed USD amount. Multiply by your SKU count and your retry rate, not by a hope.
import os
import requests
key = os.environ["SUME_API_KEY"]
r = requests.get(
"https://api.sume.com/v1/images/models/openai/gpt-image-2.5/endpoints",
headers={"Authorization": f"Bearer {key}"},
timeout=60,
)
r.raise_for_status()
body = r.json()
body = body.get("data", body)
for ep in body["endpoints"]:
print(ep["provider_slug"])
for line in ep["pricing"]:
print(" ", line["billable"], line["unit"], line["cost_usd"])Is a still cheaper than a clip?
We do not publish a single comparison number, because the two are metered differently: a still by output tokens at your quality and size, a clip by model, resolution and seconds. What the docs do let you do is compare on your own account: read the endpoint pricing for the still, run one clip and read its usage.cost, and put the two next to each other. A reasonable plan for a catalog is stills for every SKU, and a clip only for the products where motion sells the garment.
If you want a clip, the order matters. Approve the still first, then animate it, so you do not pay for motion on a garment that was wrong in frame one.
How do you keep the bill predictable?
Test at a lower quality, write your prompt once, then raise quality only for the final image. Reuse the person photo, resize it to what the model needs, and send one garment image per call. For a catalog, run a small first batch and read the real usage.cost on each response. For video, the cost shape is different: see Video generation for the models and lengths, and read each model's pricing from its catalog entry rather than from a blog.
- Set
qualityyourself;autoreserves the highest price. - Do not retry a completed generation you simply dislike without changing the prompt.
- Send an
Idempotency-Keyon each call so a network retry does not double-bill. - Read
usage.coston real responses before estimating a catalog. - Budget input tokens too: a try-on call sends two or more images.
Sources
Related posts
More in Pricing
- AI voice agent cost per minute: speech-to-text, LLM and TTS stack
A dated cost sheet for the October 2026 voice stack: MAI-Transcribe-2-Streaming, Mercury Voice, MAI-Voice-2.1-Flash, and Sume's file STT and TTS, per minute.
- Change a person in a video with AI: Sume Recast price, not free
Changing a person in a video with Sume's H3 Max Recast is paid per source second: $0.375 at 768p, $0.5625 at 1080p. Costs for 5 to 30 second clips.
- Cost per row of a Sume bulk run: add up debited_usd_micros
Read usage.debited_usd_micros on each child receipt, wait for final to be true, and treat null as unknown. Why billable_amount alone understates a bulk run.
- Cost to remove backgrounds and upscale 500 product photos by API
Sume bills background removal at $0.0225 an image and upscale at $0.20, both flat. 500 photos through both steps cost $111.25; the table and a submit loop.
Written by Sume