AI virtual try-on cost per image: GPT Image 2.5 on Sume

What one try-on still costs on Sume: the documented token rates, why quality auto reserves the max price, and a script that reads the live endpoint pricing.

5 min readSume
All posts

A try-on still on Sume is one openai/gpt-image-2.5 call, and its price is what the endpoint's pricing lines say at the quality you choose, not a flat per-try-on fee. Sume's docs give the underlying rates: $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens, and at 1024 by 1024 an xhigh output is $0.09366 and max is $0.21072 before input tokens and Sume pricing. A try-on call sends at least two input images, so input tokens count too.

ChatGPT Try On, launched October 1 2026 as a shopping feature, is not priced per image in anything we read. This post is about the API route.

Where do those numbers come from?

They are in the Image API docs, which say Flare and Sunburst use the same Fal token rates and link to the Fal model page. The docs add that output estimates use OpenAI's ChatGPT Image 2.5 size and quality calculator, that input token counts are estimates, and that Fal rounds the total up to $0.0001. We did not open the Fal page, so the rates here are exactly as Sume's docs state them.

What does a try-on call add to the output price?

Three things. First, input images: a try-on sends a person photo and a garment photo, and input image tokens are billed at the input rate. Second, Sume's margin: the docs say an endpoint's pricing lines are the amount charged to your wallet, with the margin already applied. Third, reservation: auto quality reserves max, and auto size and named presets without a verified pixel mapping reserve the output token upper bound. Omitted quality defaults to high.

What sets the price of one try-on still, read 2026-10-03
FactorDocumented ruleWhat to do
Output rate$30 per million output image tokensPick quality deliberately
Input images$8 per million input image tokens; counts are estimatesSend only the references the edit needs
1024x1024 xhigh output$0.09366 before input and Sume pricingUse as a floor for a square test
1024x1024 max output$0.21072 before input and Sume pricingAvoid unless a print needs it
quality: autoReserves the max priceSet quality explicitly
Failed generationNot billed; a completed one is billed in fullCheck results before looping

How do you read the live price before you batch?

Ask the endpoint. GET /v1/images/models/{model_id}/endpoints returns the billable lines for the model, per the docs example. The script below prints them; run it before you queue a catalog, and again if the price might have changed. The response example in the docs has no data wrapper, so the script accepts either shape.

Then run one real try-on call and read usage.cost from the response, which the docs define as the billed USD amount. Multiply by your SKU count and your retry rate, not by a hope.

import os
import requests

key = os.environ["SUME_API_KEY"]
r = requests.get(
    "https://api.sume.com/v1/images/models/openai/gpt-image-2.5/endpoints",
    headers={"Authorization": f"Bearer {key}"},
    timeout=60,
)
r.raise_for_status()
body = r.json()
body = body.get("data", body)
for ep in body["endpoints"]:
    print(ep["provider_slug"])
    for line in ep["pricing"]:
        print(" ", line["billable"], line["unit"], line["cost_usd"])

Is a still cheaper than a clip?

We do not publish a single comparison number, because the two are metered differently: a still by output tokens at your quality and size, a clip by model, resolution and seconds. What the docs do let you do is compare on your own account: read the endpoint pricing for the still, run one clip and read its usage.cost, and put the two next to each other. A reasonable plan for a catalog is stills for every SKU, and a clip only for the products where motion sells the garment.

If you want a clip, the order matters. Approve the still first, then animate it, so you do not pay for motion on a garment that was wrong in frame one.

How do you keep the bill predictable?

Test at a lower quality, write your prompt once, then raise quality only for the final image. Reuse the person photo, resize it to what the model needs, and send one garment image per call. For a catalog, run a small first batch and read the real usage.cost on each response. For video, the cost shape is different: see Video generation for the models and lengths, and read each model's pricing from its catalog entry rather than from a blog.

  • Set quality yourself; auto reserves the highest price.
  • Do not retry a completed generation you simply dislike without changing the prompt.
  • Send an Idempotency-Key on each call so a network retry does not double-bill.
  • Read usage.cost on real responses before estimating a catalog.
  • Budget input tokens too: a try-on call sends two or more images.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume