GPT Image 2.5 with no image_size quotes 3.4x the 1024 price: set one
If you omit image_size on GPT Image 2.5, Sume quotes the upper bound: $0.2225 at high against $0.0659 for 1024x1024, 3.4x. Set the size as pixels. Tier table.

If a GPT Image 2.5 request on Sume leaves out image_size, the quote reserves the upper bound of output tokens, and that is 3.4 times the price of a 1024x1024 image at every tier. At high it is $0.2225 against $0.0659, and at medium it is $0.0556 against $0.0165. Setting the size as pixels removes the gap.
Sume's Image API page says the same thing in one line: auto quality reserves max, and auto size and named presets without a verified pixel mapping reserve the upper bound. This post puts numbers on that sentence.
What the reserve costs
The reserve assumes the largest output the estimator allows, which is a 96-grid image on an 8,294,400-pixel frame. The table compares each tier with and without a size.
| Quality | 1024x1024 | No image_size | Multiple |
|---|---|---|---|
| low | $0.0074 | $0.0248 | 3.4x |
| medium | $0.0165 | $0.0556 | 3.4x |
| high | $0.0659 | $0.2225 | 3.4x |
| xhigh | $0.1171 | $0.3954 | 3.4x |
| max | $0.2635 | $0.8895 | 3.4x |
Is it billed or only reserved
Sume's docs describe the figure as a reserved bound, and usage.cost in the response is what Sume bills to your wallet. Do not assume the reserve is refunded without checking: run one call with a size and one without, and compare the two usage.cost values in your own account.
Whatever the final bill, an unset size makes any budget check based on the quote pessimistic by a factor of 3.4, and it can push a call past a spending cap that a sized call would stay under.
The named presets that do have a mapping are 1:1, 4:3, 3:4, 4:5, 5:4, 9:8, 2k and 4k, plus square_hd, landscape_4_3 and portrait_4_3. The code that prices them lives in Sume's provider-pricing package. Even so, a pixel string such as 1536x1024 leaves no room for a surprise, so prefer it to any name.
Ways to hit it by accident
The reserve is easy to trigger because several reasonable-looking requests do it.
- No
image_sizeat all, because the SDK or wrapper you use has no default. image_size: "auto", which is valid and lets the model choose the size.- A name without a verified pixel mapping, such as
16:9,9:16,landscape_16_9orsquare. These reserve the upper bound. Use pixels. - An
aspect_ratioin place of a size on a client that forwards it as the size input. quality: "auto", which reservesmaxand stacks with the size reserve.
The fix
Pass the size as pixels with both edges a multiple of 16, a longest edge up to 3840, a ratio up to 3:1 and 655,360 to 8,294,400 pixels, and set quality explicitly.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Product photo of a steel water bottle on slate",
"quality": "medium",
"image_size": "1024x1024"
}'Guard it in code
Add a check in the one place that builds the request body: if image_size is missing or auto, raise an error, unless the caller opted in. A single guard like that is cheaper than reading invoices.
Also set quality in the same guard. A missing quality means high on Sume, which is a deliberate default, but it is the second-most common reason a batch costs more than planned.
For a team, put the rule in a shared request builder and a test. The test posts a body without a size to the builder and asserts that it raises. That turns a pricing surprise into a failing build, which is the cheapest place to catch it. Add a second assertion that quality is one of the five tiers you allow, and a third that the pixel string passes the multiple-of-16 rule, so every limit in the section above is enforced before a request leaves your code. Keep these checks next to the call so the next person who edits it reads them.
Sources
Related posts
More in Pricing
- GPT Image 2 medium costs 4x GPT Image 2.5 medium; low is the same
On Sume, GPT Image 2 costs $0.0664 at medium and $0.2639 at high, about 4x GPT Image 2.5. At low the two are within 3%. Tier table at 1024x1024.
- Estimating a gpt-realtime-2.1 call from $32 / $64 per M audio tokens
gpt-realtime-2.1 lists audio input at $32, cached input at $0.40 and audio output at $64 per million tokens. Here is the formula and a worked example.
- gpt-transcribe at $0.0045 a minute vs Scribe v2 and Nova-3
Of the three, Scribe v2 lists lowest at $0.22 an hour, then Nova-3 mono at $0.258 and gpt-transcribe at $0.27. Per-hour table and what Sume bills.
- Grok Imagine image prices: $0.02, $0.04, $0.05 and Sume's one row
xAI lists three Grok image prices ($0.02, $0.04, $0.05). Sume lists one x-ai/grok-image row; read its live endpoints pricing array for the billed figure.
Written by Sume