GPT Image 2.5 at 1024px: xhigh is 3,122 output tokens, max is 7,024
Sume's docs give GPT Image 2.5 output estimates at 1024x1024: xhigh $0.09366, max $0.21072 at $30 per million tokens, i.e. 3,122 and 7,024 tokens.

Sume's Image API docs give two output estimates for GPT Image 2.5 at 1024 x 1024: $0.09366 at xhigh and $0.21072 at max. At $30 per million output tokens those are 3,122 and 7,024 tokens, so max produces 2.25 times the output tokens of xhigh. Both figures are before input tokens and before Sume's margin.
Converting dollars to tokens
The docs state that Flare and Sunburst use the same fal token rates: $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens. OpenAI's own pricing page lists the same $30.00, $8.00 and $5.00 Standard rates for GPT Image 2.5 (read 2026-10-09).
| Quality | Output cost | Arithmetic | Output tokens |
|---|---|---|---|
| xhigh | $0.09366 | 0.09366 / 0.00003 | 3,122 |
| max | $0.21072 | 0.21072 / 0.00003 | 7,024 |
| Ratio max to xhigh | 2.25 | 7,024 / 3,122 | 2.25 |
With Sume's margin
Multiplying by 1.25 gives $0.117075 for xhigh and $0.2634 for max, output only. Input tokens add to that, and the docs say input counts are estimates and that fal rounds the total up to $0.0001.
Sume's fixed catalog rows are lower because they cover the lower tiers: low at 1K is $0.02475, medium at 2K is $0.055625 and high at 4K is $0.2225. A 4K high image at $0.2225 is below the 1024 max output estimate with margin, so max is not a quality step to take casually.
Choosing a quality
For 500 images the output-only gap between the two top tiers at 1024 is 500 x ($0.21072 - $0.09366) = $58.53 before margin. Unless you need max, xhigh or high is the cheaper way to get dense text and fine detail. Set the quality explicitly because auto reserves max.
Run a small sample of 10 at each tier on your own prompt, compare the outputs, and pick the cheapest tier whose output you would ship.
Caveats
These are output-token estimates for one size, 1024 x 1024. Other sizes produce other token counts, and Sume's docs say named presets without a verified GPT-specific pixel mapping reserve the upper bound of output tokens. Input counts are estimates too.
Treat the arithmetic as a way to understand the scale, and use the endpoint's pricing lines for the amount your wallet is actually charged.
Sources
Related posts
More in Models
- Grok Imagine Video 1.5 on Sume: image-in only, silent, $0.19 for 15 s
Grok Imagine Video 1.5 on Sume needs a start image, makes silent 480p or 720p clips of 4 to 15 seconds, and a 15-second clip is 15 x $0.0125 = $0.1875.
- Haiku 5.5 cache write vs read: break-even on a Sume tool catalog
Haiku 5.5 charges $0.125 per million to write cache and $0.01 to read it. For a 12,000-token Sume tool list the cache pays back on the first reuse.
- HappyHorse 1.1 image-to-video: 300 px, 20 MB and ratio 1:2.5 to 2.5:1
Alibaba's HappyHorse 1.1 image-to-video accepts JPEG, PNG or WEBP of at least 300 px, up to 20 MB. Prepare a first frame, or use an Omni image_url on Sume.
- HappyHorse 1.1 API watermark defaults to on: how to turn it off
Alibaba's HappyHorse 1.1 image-to-video API adds a Happy Horse text mark bottom-right unless you set watermark to false. What to check before delivery.
Written by Sume