GPT Image 2.5 cached input works only on the Responses API
OpenAI's cached-input rates ($1.25-$2.00 per 1M) for GPT Image 2.5 apply only to Responses API images. Sume bills token rates with margin and no cache discount.

OpenAI's cached-input rates for GPT Image 2 and GPT Image 2.5 apply only to images generated with the Responses API, so a single-request Image API workflow gets no caching benefit. Multi-turn editing, where you resend the same context, is the case that can use it. Sume's ChatGPT Image 2.5 rows do not list a cache discount.
OpenAI's rates
Rates are per 1M tokens from OpenAI's pricing page. Batch is half.
| Token type | Price per 1M tokens |
|---|---|
| Text input | $5.00 |
| Image input | $8.00 |
| Cached input | $1.25 to $2.00 |
| Image output | $30.00 |
Which API fits which workflow
OpenAI's guide describes the Image API as a single request and the Responses API as the multi-turn route. Cached input only helps when the same input tokens come back across turns, which is the Responses pattern.
| Workflow | API | Cached input |
|---|---|---|
| One prompt, one image | Image API | No |
| Iterating on one image over several turns | Responses API | Yes, per OpenAI's pricing note |
| Batch overnight | Either, at batch (half) rates | Check the pricing page for your combination |
On Sume
Sume's Flare and Sunburst bill token rates of $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens at list. Sume adds its margin (list x 1.25) and returns the billed USD in usage.cost; token counts in the response are 0. Each call to POST /v1/images is a separate request, so there is no cache to hit.
How to choose
If you iterate on one image many times and cost matters, price that on OpenAI's Responses API. If you want one billing line and one schema across many image models, use Sume. The right choice depends on how many turns you run, which only your workload can say.
Sources
Related posts
More in Developers
- GPT Image 2.5 output_compression returns 400 on Sume: use Pillow
OpenAI offers output_compression 0-100 for JPEG and WebP; Sume returns 400 unsupported_parameter for it. Request png, jpeg or webp, then compress with Pillow.
- Graph API rate-limit codes 4, 17, 32, 613: stop, reuse the Sume job
Meta documents error codes 4, 17, 32 and 613 for rate limits. Map each to a pause and publish the stored Sume result later instead of re-rendering.
- Graph API X-App-Usage call_count: pause a Reel publisher early
Meta's X-App-Usage header reports call_count, total_cputime and total_time as percentages. Read it after each call and pause your publisher before the limit.
- Griffin 10 ms audio packets vs Sume TTS: async jobs, no streaming
Tavus Griffin emits speech in packets as small as 10 ms. Sume's TTS Router lists streaming TTS as a non-goal, so audio arrives as a finished file for lip sync.
Written by Sume