GPT Image 2.5 at 1024px: xhigh is 3,122 output tokens, max is 7,024

Sume's docs give GPT Image 2.5 output estimates at 1024x1024: xhigh $0.09366, max $0.21072 at $30 per million tokens, i.e. 3,122 and 7,024 tokens.

5 min readSume
All posts

Sume's Image API docs give two output estimates for GPT Image 2.5 at 1024 x 1024: $0.09366 at xhigh and $0.21072 at max. At $30 per million output tokens those are 3,122 and 7,024 tokens, so max produces 2.25 times the output tokens of xhigh. Both figures are before input tokens and before Sume's margin.

Converting dollars to tokens

The docs state that Flare and Sunburst use the same fal token rates: $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens. OpenAI's own pricing page lists the same $30.00, $8.00 and $5.00 Standard rates for GPT Image 2.5 (read 2026-10-09).

GPT Image 2.5 output at 1024x1024, from Sume docs, as of 2026-10-09
QualityOutput costArithmeticOutput tokens
xhigh$0.093660.09366 / 0.000033,122
max$0.210720.21072 / 0.000037,024
Ratio max to xhigh2.257,024 / 3,1222.25

With Sume's margin

Multiplying by 1.25 gives $0.117075 for xhigh and $0.2634 for max, output only. Input tokens add to that, and the docs say input counts are estimates and that fal rounds the total up to $0.0001.

Sume's fixed catalog rows are lower because they cover the lower tiers: low at 1K is $0.02475, medium at 2K is $0.055625 and high at 4K is $0.2225. A 4K high image at $0.2225 is below the 1024 max output estimate with margin, so max is not a quality step to take casually.

Choosing a quality

For 500 images the output-only gap between the two top tiers at 1024 is 500 x ($0.21072 - $0.09366) = $58.53 before margin. Unless you need max, xhigh or high is the cheaper way to get dense text and fine detail. Set the quality explicitly because auto reserves max.

Run a small sample of 10 at each tier on your own prompt, compare the outputs, and pick the cheapest tier whose output you would ship.

Caveats

These are output-token estimates for one size, 1024 x 1024. Other sizes produce other token counts, and Sume's docs say named presets without a verified GPT-specific pixel mapping reserve the upper bound of output tokens. Input counts are estimates too.

Treat the arithmetic as a way to understand the scale, and use the endpoint's pricing lines for the amount your wallet is actually charged.

Sources

Related posts

More in Models

All Models posts

Written by Sume