GPT Image 2.5 image_size auto: priced as 8,294,400 pixels
Sume's docs say image_size auto on GPT Image 2.5 reserves the upper bound of output tokens. Estimate: $0.2224 at high vs $0.0659 at 1024x1024. Send a size.
On Sume, leaving GPT Image 2.5 on auto size is priced against the upper bound: 8,294,400 pixels, the largest image the model accepts. By my estimate that is about $0.2224 at high quality, against about $0.0659 for a 1024x1024 image, or roughly 3.4 times as much. If you know the size you need, send it.
The estimates below assume a text prompt with no reference images, and every figure is before cent rounding.
What the docs say
Sume's Image API docs state that auto quality reserves max, and that auto size, along with named presets that lack a verified GPT-specific pixel mapping, reserve the upper bound of output tokens. OpenAI's image guide, read on 2026-10-07, gives that bound: sizes may total between 655,360 and 8,294,400 pixels.
The two settings are separate. Quality auto reserves the max tier; size auto reserves the largest pixel count. A request with both on auto has the largest reservation the model can have, and the cost check uses it.
Sume's docs add that auto model routing continues to use Flare, and they do not say which size the model would pick on its own. Without that information, you cannot predict the image you will get, only the ceiling you will be reserved against, which is a second reason to name a size.
The estimates
The figures below come from my calculation using the output-token formula in Sume's pricing code, the $30 per million output token rate, and the 1.25 multiplier. They are for a text-only prompt, before cent rounding and before reference-image input tokens. The final amount is in usage.cost.
The high auto row works out as follows: the upper bound gives 5,930 output tokens, which is $0.17790 at list; times 1.25 that is $0.222375, which rounds up to $0.23 at the cent.
In ratio terms, the auto row is 3.4 times the 1024x1024 row at high ($0.2224 divided by $0.0659), 3.4 at low, and 3.4 at max, because both columns scale with the same quality grid. The conclusion does not depend on the tier you pick.
| Quality | 1024x1024 | Auto size |
|---|---|---|
| Low | $0.0074 | $0.0248 |
| Medium | $0.0165 | $0.0556 |
| High | $0.0659 | $0.2224 |
| Xhigh | $0.1171 | $0.3954 |
| Max | $0.2635 | $0.8895 |
Send a fixed size
Name the pixels. The custom image_size object takes a width and a height, both multiples of 16, with a long edge up to 3840 and a ratio up to 3:1. For a 16:9 image, 1536x864 works and stays well below the upper bound. The request below pins model, quality and size.
curl -sS -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "A minimalist poster of a mountain at sunrise",
"quality": "medium",
"image_size": {"width": 1536, "height": 864}
}'What to check in your code
Search for requests that set no image_size, or that set it to auto, and for any quality left out or set to auto. Omitting quality means high on GPT Image 2.5, not auto, so the two defaults differ. A budget guard that reads the catalog price before submitting will use the reservation, which is why a small image can look expensive in the guard.
A practical default: define two or three sizes in your configuration, such as 1024x1024, 1536x864 and 864x1536, and let callers choose by name. This keeps both the picture and the bill predictable, and it makes every size pass the multiple-of-16 rule by construction.
Sources
Related posts
More in Pricing
- GPT Image 2.5 quality ladder: price multiplier from low to max on Sume
At 1024x1024 on Sume, GPT Image 2.5 high costs 8.9 times low, xhigh 15.9 times and max 35.7 times. The ladder with billed prices and batch costs.
- GPT Image 2.5 xhigh vs max at 1024x1024: $0.09366 vs $0.21072
At 1024x1024, Sume's docs put GPT Image 2.5 xhigh output at $0.09366 and max at $0.21072 before input tokens and Sume pricing. When max is worth 2.25x.
- H3 Max lip sync at 14.8 s bills 15 s: $1.50 at 768p, $3.00 at 1080p
A 14.8-second line is the longest MiniMax H3 Max lip sync accepts on Sume. It rounds up to 15 billed seconds: $1.50 at 768p, $3.00 at 1080p, $0.9375 at 480p.
- MiniMax H3 Max API price: Sume 1.25x vs the fal list
H3 Max costs $0.0625 to $0.20 per second on Sume (list x 1.25). Per-clip and per-100-clip math against fal's listed $0.05 per second.
Written by Sume