GPT Image 2.5 quality auto reserves max: set quality before a batch
On Sume, quality auto reserves the max token bound for GPT Image 2.5. Omitting quality uses high. Pick the level yourself before you queue a batch.

On Sume, quality: "auto" for GPT Image 2.5 reserves the max bound, the most expensive level, while an omitted quality uses high. So auto is not a cheap default. Before a batch, set quality explicitly, and check the endpoint pricing line with GET /v1/images/models/openai/gpt-image-2.5/endpoints.
Three behaviors to know
The docs describe how the quality field resolves, and the three cases differ in what you pay and what you reserve.
- Omitted: Sume sends
high, and the estimate priceshightoo, so the reserve matches the run. auto: Sume reserves themaxbound. The documented auto size also reserves the upper bound of output tokens.- Explicit
low,medium,high,xhighormax: Sume prices exactly that level for the requested size.
What the levels cost at 1024 square
Flare and Sunburst share the same token rates: $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens. The docs give the output estimate at 1024 by 1024 for the two top levels. Input tokens are extra, and Sume applies its list times 1.25 pricing on top.
| quality | Output estimate |
|---|---|
| high | about $0.0527 (catalog list row, rounded to $0.0001) |
| xhigh | $0.09366 |
| max | $0.21072 |
Why this matters for a batch
A balance check runs against the reserve. If you queue 200 jobs with quality: auto, each is admitted at the max bound, so a wallet that comfortably covers 200 high images can still be short at admission. The final bill follows the completed generation, but the hold is what blocks the queue. Setting the level yourself keeps the hold close to the real cost.
A safe request
Pin the level and the size in the body. Use medium for drafts, high for finals, and keep xhigh and max for dense text or print work.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-image-2.5","prompt":"Clean product photo of a steel water bottle on a stone ledge, soft morning light","quality":"medium","aspect_ratio":"4:5"}'Check the price you will actually pay
The endpoint pricing lines are the amount Sume charges your wallet, with margin included, so you pay cost_usd times n. Read it once per model and cache it. If the line changes, your budget guard should notice. See the GPT Image 2.5 API guide for sizes and masks.
Related posts
More in Pricing
- GPT Image 2.5 sizes 1024x1024, 1536x1024, 1024x1536: Sume price
OpenAI's three recommended GPT Image 2.5 sizes do not cost the same on Sume: 1536x1024 high is $0.0515 against $0.0659 square. Prices by tier and size.
- GPT Image 2 medium costs 4x GPT Image 2.5 medium; low is the same
On Sume, GPT Image 2 costs $0.0664 at medium and $0.2639 at high, about 4x GPT Image 2.5. At low the two are within 3%. Tier table at 1024x1024.
- Estimating a gpt-realtime-2.1 call from $32 / $64 per M audio tokens
gpt-realtime-2.1 lists audio input at $32, cached input at $0.40 and audio output at $64 per million tokens. Here is the formula and a worked example.
- gpt-transcribe at $0.0045 a minute vs Scribe v2 and Nova-3
Of the three, Scribe v2 lists lowest at $0.22 an hour, then Nova-3 mono at $0.258 and gpt-transcribe at $0.27. Per-hour table and what Sume bills.
Written by Sume