Eight low-quality edits cost less than one high GPT Image 2.5 edit
On Sume, a GPT Image 2.5 edit with one reference costs $0.0094 at low and $0.0835 at high, so eight low edits cost $0.075, less than one high edit.

A GPT Image 2.5 edit with one reference image costs $0.0094 at low, $0.0209 at medium and $0.0835 at high on Sume at 1024x1024, so eight low edits cost $0.0750, which is less than a single high edit at $0.0835. A ninth low edit, at $0.0844, would tip it the other way.
The point is the shape of the quality ladder, not the arithmetic. An edit spends most of its money on output tokens, and output tokens scale with the quality tier, so exploring at low is cheap enough to try several variations of an instruction before one render at high.
Edit prices by tier
All rows are one reference image at 1024x1024 with a short instruction. The last column is how many edits at that tier fit in the price of one high edit.
| Quality | Edit price | Text-to-image price | Edits per one high edit |
|---|---|---|---|
| low | $0.0094 | $0.0074 | 8.9 |
| medium | $0.0209 | $0.0165 | 4.0 |
| high | $0.0835 | $0.0659 | 1.0 |
A loop that uses this
Send the same reference and instruction at low with two or three phrasings, pick the best, and then send the winning instruction at high. The reference stays the same, so the input side of the bill is the same on every call; only the output quality changes.
On those numbers, a three-draft loop costs $0.0281 for the drafts plus $0.0835 for the final, or $0.1116 together. Running the same three phrasings directly at high would cost $0.2505.
The saving shrinks as the number of drafts grows. Ten low drafts cost $0.0938, already more than one high edit, so cap the draft count. Eight is the most that stays under one high edit; past that, going straight to medium at $0.0209 is the better compromise.
Keep the reference and size identical between draft and final. A different size changes the output token count and the image the model sees, so a draft at one size is only a rough guide to a final at another.
Where the shortcut fails
Cheap drafts are only useful if they predict the final.
- Fine detail. Text on a label, small faces and hands are the details a
lowrender is most likely to get wrong, so a draft can pass a composition check and still fail a detail check athigh. - Masks. If your edit uses
mask_url, test the mask atlowfirst; a bad mask wastes a call at any tier. - Failed calls. A failed generation is not billed, so retries on errors are free, but a poor result that completes is billed in full.
- More references. Each extra reference image adds about 27% of the output price to the call, so a three-reference edit at
lowcosts $0.0132, not $0.0094.
The same ladder at 1536x1024
A landscape frame follows the same pattern at lower prices. A one-reference edit at 1536x1024 costs $0.0076 at low, $0.0164 at medium and $0.0653 at high, so eight low edits cost $0.0610, still under the high price, and a ninth would pass it. The drafts-per-final count changes with the frame, so recompute it for your own size and tier instead of reusing eight.
If you edit a batch, treat the first few images as drafts and the rest as finals. Run every image at low, review the results, and re-run only the ones whose composition is right at a higher tier. The review step is where the saving comes from, since each rejected image costs under a cent at 1024x1024 and less at 1536x1024. Keep a simple tally of kept against rejected drafts, because the keep rate decides whether the loop beats going straight to high.
Request
One edit call: the reference goes in image_urls. Change quality to high for the final render.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Make the sky sunset orange, keep the building unchanged",
"image_urls": ["https://example.com/building.png"],
"quality": "low",
"image_size": "1024x1024"
}'Sources
Related posts
More in Pricing
- What Eleven v4 costs per hour of audio before the Oct 12 promo ends
At $0.022 per 1K characters, an hour of narration is about $1 on Eleven v4 during the promo. Sonic via Sume is about $2.14 to $2.28 for the same text.
- Music API cost: ElevenLabs $0.15/min vs Lyria $0.08/song
ElevenLabs lists music at $0.15 a minute; Google lists Lyria 3.5 at $0.08 a song. A two-minute track is $0.30 vs $0.08. Sume is a fixed $0.125 per generation.
- Scribe v2 at $0.22 an hour vs what Sume bills per audio minute
ElevenLabs lists Scribe v2 at $0.22 per hour (about $0.0037 a minute). Sume bills speech-to-text at $0.01 per audio minute, or $0.60 an hour. Why they differ.
- Text to speech API price per 1,000 characters in October 2026
From $0.011 (Eleven v4 Turbo, promo) to $0.08 (Eleven v3) per 1K characters, with Deepgram, OpenAI and Sume's Sonic at $0.0475 in one dated table.
Written by Sume