GPT Image 2.5 auto quality: why the reserve is the max price

quality auto on GPT Image 2.5 reserves the max amount on Sume, not the amount you end up paying. How reserve differs from charge, and how to size a batch.

4 min readSume
All posts

With quality: "auto", Sume reserves the max amount for the request, which is $0.27 for one 1024×1024 image against $0.07 for high. The reserve is a hold sized for the worst case, not the final charge. The docs say actual cost is captured on success, so an auto image that lands cheaper should settle for less than it held.

Sources are Sume's Image API docs, Core concepts and the pricing code, read 2026-09-29. The usage.cost in each response is the USD amount billed.

What does the reserve look like next to the default?

Leaving quality out runs high, and the reserve matches that tier. Choosing auto moves the reserve to the top of the ladder even if the picture needs less.

Reserve for one 1024x1024 GPT Image 2.5 image, Sume estimates from pricing code, read 2026-09-29. Output tokens only.
You sendReserved on submit
Nothing (runs high)$0.07
quality: "auto"$0.27 (the max amount)
Ten requests at auto$2.70

Why can auto trigger a 402 with money in the wallet?

Because the reserve, not the eventual charge, has to fit. The API returns 402 insufficient_credits when the balance is not sufficient for the requested generation, and provider-backed generation can reserve its estimate up front. A balance that easily covers ten high images may fall short of ten auto reserves held at the same moment.

Does the size setting change the reserve too?

Yes. auto size, and named presets without a verified pixel mapping, reserve the output-token upper bound. Set an explicit image_size such as 1024x1024 and an explicit quality, and the reserve tracks the request you meant to make.

curl -X POST https://api.sume.com/v1/images \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-image-2.5",
    "prompt": "A red bicycle against a blue wall",
    "image_size": "1024x1024",
    "quality": "medium"
  }'

When is auto still the right choice?

When you would rather let the provider choose the effort and can carry the larger hold. Keep a balance above the sum of the reserves you plan to have in flight, then read usage.cost on the responses to see what each image really cost. A failed or cancelled generation before capture is refunded.

How do I size the wallet for a batch?

Multiply the reserve by the number of requests you will have in flight at once, not by the total in the batch. Reserves are released or settled as each generation finishes, so a long queue run five at a time holds five reserves, not five hundred. Pin quality and image_size and that number gets much smaller.

Then compare with the response: usage.cost is what you were actually billed. If your auto images settle well under the reserve, you can lower the concurrency need by setting the tier you actually get. If they settle near the reserve, auto is simply choosing high effort and you should say so in the request.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume