gpt-image-2.5 quality auto reserves max: holds from $0.22 to $0.89
On Sume, gpt-image-2.5 with quality auto and auto size reserves $0.8895 per image, while omitting quality reserves $0.2224. Hold table and the safe request.

Sending quality: "auto" with no size reserves $0.8895 per gpt-image-2.5 image on Sume, compared with $0.2224 when you omit quality and $0.0659 when you also pin 1024x1024. The Image API states that auto quality reserves max, and that auto size reserves the upper bound of output tokens.
The reserved amount is a hold at submit, not necessarily the final charge, but a hold you cannot cover returns 402 insufficient_credits before work starts.
Hold by setting
| Quality | Size | Reserved per image |
|---|---|---|
| omitted (high) | auto | $0.2224 |
| auto | auto | $0.8895 |
| auto | 1024x1024 | $0.2635 |
| high | 1024x1024 | $0.0659 |
| medium | 1024x1024 | $0.0165 |
| low | 1024x1024 | $0.0074 |
The safe request
Pin both. A low balance then fails because of a number you chose.
import os, requests
r = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={
"model": "openai/gpt-image-2.5",
"prompt": "Flat lay of three ceramic bowls",
"quality": "medium",
"image_size": {"width": 1024, "height": 1024},
},
timeout=60,
)
print(r.status_code)
body = r.json()
if r.status_code == 200:
for item in body["data"]:
print(item["url"])
else:
print(body)Why it matters for batches
A loop of 1,000 requests that leaves both on auto needs $889.50 of headroom per thousand, against $16.50 when pinned to medium. Concurrent requests each hold their own reservation, so unpinned batches can hit 402 far earlier than the real spend suggests.
Rule of thumb
- Always set
quality. - Always set
image_sizeor a named size you have priced. - Treat the table above as the hold your balance must cover per in-flight request.
- Read
GET /v1/images/models/{model_id}/endpointsfor the livepricingline before a large run: prices are catalog data and can change.
Sources
Related posts
More in Developers
- Haiku 5.5 prompt caching for a Sume tool list: what breaks the cache
Haiku 5.5 cache hits cost $0.01 per million tokens. Keep the Sume tool list and effort setting stable so a long agent run keeps hitting the cache.
- Haiku 5.5 returns 400 for thinking disabled at xhigh: a Sume agent fix
Claude Haiku 5.5 rejects thinking disabled at xhigh or max effort. How that 400 shows up in an agent that calls Sume, and how to set the pair.
- Hono 4.13.10 split adapters: update a Sume webhook receiver
Hono 4.13.10 moved adapters to @hono/bun, @hono/deno and @hono/cloudflare-workers. The Sume verifyWebhook call needs no change; only the entrypoint imports do.
- Hono 4.13.13 deprecates app.mount(): mount a Sume webhook sub-app
Hono 4.13.13 deprecates app.mount() in favor of Mount Middleware. A Sume webhook receiver written as a Hono sub-app uses app.route() and keeps its raw body.
Written by Sume