GPT Image 1 retires Oct 23: a 20-prompt test on 2.5 for about $6.06
OpenAI lists gpt-image-1 for removal on October 23, 2026. Test 20 prompts at low, medium and high on gpt-image-2.5 through Sume for about $6.06 first.

OpenAI's deprecations page lists gpt-image-1 with a shutdown date of October 23, 2026 and names gpt-image-2.5-sunburst or gpt-image-2.5-flare as the recommended replacement (read 2026-10-09). That leaves 14 days from today. A 20-prompt regression test on Sume's gpt-image-2.5 row at all three quality tiers costs about $6.06, using the catalog prices as of 2026-10-09.
The point of the test is to find out which of your prompts change when the model changes, before a customer finds out for you.
What the test costs
Sume bills gpt-image-2.5 by quality and size: $0.02475 at low and 1K, $0.055625 at medium and 2K, and $0.2225 at high and 4K. Twenty prompts at each tier is 60 images.
| Tier | Per image | 20 images |
|---|---|---|
| low, 1K | $0.02475 | $0.49 |
| medium, 2K | $0.055625 | $1.11 |
| high, 4K | $0.2225 | $4.45 |
| All three tiers | $0.302875 | $6.0575 (about $6.06) |
How to run it
Take 20 prompts from your real traffic, not 20 new ones. Pick them to cover the failure modes you care about: text in the image, a product on white, a person, a long prompt, a prompt with reference images. Save the old gpt-image-1 outputs if you still have them, since you cannot make new ones after October 23.
Run each prompt at low first. Low is the cheapest tier at 2.5 cents, and it shows prompt-following problems well. Run medium and high only on the prompts where low looks wrong or where your product uses a higher tier.
- Keep one seed-free, unchanged prompt per case. Sume does not serve a
seedfield, so two runs of the same prompt will differ. - Record
usage.costfrom each response so the test bill matches the numbers above. - Expect some 4K and high calls to answer with a 202 job envelope. Poll the job and fetch the result.
Why not just wait
After October 23 the old id stops working, so a test you run later has nothing to compare against. A two-hour test now costs less than one hour of an engineer. If your product uses gpt-image-1 only for a few fixed prompts, the test can be shorter: ten prompts at low cost 25 cents.
What to compare
Check four things per case: whether the requested text is spelled correctly, whether the aspect ratio is honored, whether reference images were used, and whether the file format is what your pipeline expects. If a case fails on 2.5, change one thing at a time in the prompt and rerun at low, which costs 2.5 cents per try.
OpenAI also lists chatgpt-image-latest, gpt-image-1-mini and gpt-image-1.5 for removal on December 1, 2026. If your code uses any of those ids, put them on the same test plan. The Image API docs list the model ids that Sume accepts.
On Sume you send the model id in the request body, so switching is a one-line change. Check your own client for hard-coded ids in config files and in tests.
Sources
Related posts
More in Models
- GPT Image 2.5 at 1024px: xhigh is 3,122 output tokens, max is 7,024
Sume's docs give GPT Image 2.5 output estimates at 1024x1024: xhigh $0.09366, max $0.21072 at $30 per million tokens, i.e. 3,122 and 7,024 tokens.
- GPT Image 2.5 quality auto reserves max: set quality yourself on Sume
On Sume, GPT Image 2.5 quality auto reserves max. At 1024px the max output estimate is $0.2634 with margin versus a $0.02475 low row. Set quality explicitly.
- Grok Imagine Video 1.5 on Sume: image-in only, silent, $0.19 for 15 s
Grok Imagine Video 1.5 on Sume needs a start image, makes silent 480p or 720p clips of 4 to 15 seconds, and a 15-second clip is 15 x $0.0125 = $0.1875.
- Haiku 5.5 cache write vs read: break-even on a Sume tool catalog
Haiku 5.5 charges $0.125 per million to write cache and $0.01 to read it. For a 12,000-token Sume tool list the cache pays back on the first reuse.
Written by Sume