GPT Image 2.5 on Sume: each reference adds about 27% to the quote
Each reference image adds about 27% of the output price to a GPT Image 2.5 call on Sume: $0.0165 at medium rises to $0.0209 with one and $0.0868 with 16.

Each reference image you add to a GPT Image 2.5 request raises the Sume quote by about 27% of the output price: a medium 1024x1024 call goes from $0.0165 with no reference to $0.0209 with one and $0.0867 with sixteen. The step is $0.0044 for the first reference at that tier, and the later ones add about the same.
The reason is in Sume's estimator. Sume's Image API page says input image tokens are billed at $8 per million, against $30 per million for output image tokens, and that the counts are estimates. The estimator assumes one reference is about as many tokens as the output image, so each reference costs 8/30 of the output price, or 26.7%, before Sume's 1.25 factor.
Price by reference count
All rows are medium quality at 1024x1024 with a short prompt. The second column is Sume's quote, and the third is the change from the text-to-image price.
| References | Sume price | Change from zero |
|---|---|---|
| 0 | $0.0165 | +0% |
| 1 | $0.0209 | +27% |
| 2 | $0.0253 | +53% |
| 4 | $0.0341 | +107% |
| 8 | $0.0516 | +213% |
| 16 | $0.0867 | +426% |
The same rule at other tiers
Because the reference cost is a share of the output cost, the percentage is about the same at every tier and the dollars scale with the tier. One reference at low takes the quote from $0.0074 to $0.0094, a rise of 27%, and at high from $0.0659 to $0.0835, also 27%. Three references at low cost $0.0132, which is 80% more than text-to-image.
That makes references a larger share of the bill at the low tiers. At low three references nearly double the price. At high three references add the same 80%, but the dollars are larger: $0.0527 more per call.
How to keep the cost down
Reference images are often what makes an edit work, so cut them only where they do not.
- Send only the references the instruction uses. Sixteen is the limit, but a call with four costs $0.0341 at
medium, less than half of the 16-reference price. - Draft with fewer references at
low, then add the extra ones for the final render. - Use
mask_urlto restrict an edit, not extra references that describe the same region. - Check
usage.coston the first response of a new workflow against this table, because the input counts are estimates.
Planning a batch
For a batch of 1,000 edits at medium with two references each, the quote is $25.25, against $16.50 for the same 1,000 images with no reference. The references account for $8.75 of that. If the same two references are reused on every call, the cost still repeats on every call, because each request carries its own input tokens.
This also changes the decision between one call with n: 4 and four separate calls. Four images in one call are priced as four images, and the references are counted once per output image in the estimator, so a single call with four outputs does not make the references free. Compare a real usage.cost before you assume a saving, and do it once per workflow.
Finally, treat the 27% as a rule of thumb for budgeting, not as a rate card. The estimator is allowed to change its input-token assumption, and the live usage.cost and the catalog are the sources of record, and a quick re-measure each quarter is cheap. Record the date and the figure next to the number in your own budget sheet so a later reader knows how old it is.
Request
Two references in image_urls, medium quality. The response should quote close to $0.0253.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Put the jacket from the first image on the person in the second",
"image_urls": [
"https://example.com/jacket.png",
"https://example.com/person.png"
],
"quality": "medium",
"image_size": "1024x1024"
}'Sources
Related posts
More in Pricing
- GPT Image 2.5 with no image_size quotes 3.4x the 1024 price: set one
If you omit image_size on GPT Image 2.5, Sume quotes the upper bound: $0.2225 at high against $0.0659 for 1024x1024, 3.4x. Set the size as pixels. Tier table.
- GPT Image 2.5 sizes 1024x1024, 1536x1024, 1024x1536: Sume price
OpenAI's three recommended GPT Image 2.5 sizes do not cost the same on Sume: 1536x1024 high is $0.0515 against $0.0659 square. Prices by tier and size.
- GPT Image 2 medium costs 4x GPT Image 2.5 medium; low is the same
On Sume, GPT Image 2 costs $0.0664 at medium and $0.2639 at high, about 4x GPT Image 2.5. At low the two are within 3%. Tier table at 1024x1024.
- Estimating a gpt-realtime-2.1 call from $32 / $64 per M audio tokens
gpt-realtime-2.1 lists audio input at $32, cached input at $0.40 and audio output at $64 per million tokens. Here is the formula and a worked example.
Written by Sume