ChatGPT Image 2 vs 2.5 on Sume: $0.26375 vs $0.065875 and what differs
On Sume, GPT Image 2.5 high quality at 1024 costs $0.065875, a quarter of GPT Image 2 at $0.26375, and adds mask_url, background and 16 references.

On Sume, GPT Image 2.5 at high quality and 1024 square costs $0.065875, about a quarter of GPT Image 2 at $0.26375, and the newer row lists more: 16 references against 10, a mask_url, a background field and six quality levels against three. So for most new work the cheaper and newer row is the default, and sume/auto already resolves to it.
These figures come from the catalog code and the Image API docs, read on 2026-10-10. The fal page for the GPT Image 2.5 Flare variant, read the same day, lists per-quality prices for 1024 square that match the ladder in the table below: $0.0059, $0.0132, $0.0527, $0.0937 and $0.2107 list.
Contract differences
GPT Image 2 lists nine ratios including auto, 10 references and quality of low, medium or high. GPT Image 2.5 lists nine ratios including auto, 16 references, six quality levels, background and mask_url. Sume's docs say transparent generated images are possible only through the GPT Image 2.5 background field.
| Setting | GPT Image 2 | GPT Image 2.5 |
|---|---|---|
| Billed at default quality, 1024 | $0.26375 | $0.065875 |
| References | up to 10 | up to 16 |
| Quality levels | 3 | 6 |
| mask_url | no | yes |
| background (transparent) | no | yes |
| Metering | flat plus $0.013 per input image | tokens, edits add input tokens |
Pricing behaves differently
GPT Image 2 is a flat price per output image plus $0.013 per input image. GPT Image 2.5 is token metered, so the price depends on quality and size, and an edit adds an estimate for its input tokens. The $0.065875 figure is for high quality at 1024, so a lower quality or smaller size costs less and a larger one costs more.
This means you should price your own mix. Ask the endpoints route for the row, read the pricing lines, and compute the cost of a typical edit rather than assuming the headline number.
When to keep GPT Image 2
The only reason to stay on the older row is output you have already approved and want to reproduce. Pin openai/gpt-image-2 by id, and note that a pinned id is the only way to get repeatable results: the alias sume/auto resolves to GPT Image 2.5 and hides the family.
For a new brand asset pipeline, test the 2.5 row first at a lower quality and move up only when the output needs it.
- Use 2.5 for masks, transparency and many references.
- Pin ids when you need repeatability.
- Check the wallet charge as
cost_usdtimesn.
What to test before switching
Run your five most common prompts on both rows at matching quality. Compare text accuracy, edge handling and how closely each follows a reference. The catalog tells you what each row accepts; it cannot tell you which looks better for your brand, so decide with your own samples.
Switching a production pipeline is a pin change: replace openai/gpt-image-2 with openai/gpt-image-2.5 in config, keep the old id for rollback and log both in your results.
Cost at volume
At the headline prices, 1,000 images cost $263.75 on GPT Image 2 and $65.875 on GPT Image 2.5 at high quality and 1024. The difference of $197.875 per thousand is why the migration is worth a test. A lower quality on 2.5 widens the gap further, because token metering scales with quality.
Features you gain on 2.5
The mask_url field lets you confine an edit to a region, and the background field is the one route in the Sume catalog to a transparent generated image. Sixteen references leave more room for a style sheet plus product shots. The six quality levels give finer control over cost than the three on the older row.
None of those fields exist on the older row, so a request that sends them to openai/gpt-image-2 returns a 400.
If you use the Sunburst variant of 2.5, which Sume also lists, the catalog shows the same 16-reference limit and the same billed price as the main 2.5 row, so the same arithmetic applies to it.
Sources
Related posts
More in Comparisons
- Creatomate RenderScript vs a Sume Timeline document: field map
Creatomate's RenderScript is a general scene JSON; Sume's Timeline 1.0 is one audio spine plus video slots. Field-by-field map and what Sume cannot express.
- Creatomate template modifications vs a Sume Format run input
Creatomate fills a fixed template through modifications; a Sume Format run takes free-form input and generates new media. Request shapes and when each fits.
- D-ID audio talks: 15 MB, 5-10 minutes vs Sume's 4-60 s script
D-ID takes up to 40,000 characters of text or 15 MB of audio for a talk. Sume takes an English script of 4 to 60 seconds. Which input fits (read 2026-10-10).
- ElevenLabs key scopes, quota and IP allowlist vs Sume key controls
ElevenLabs keys can be scoped, capped and IP-locked. Sume keys fix scopes at creation and add per-run spend caps. What each covers and what Sume docs do not.
Written by Sume