GPT Image 2.5 drafts at low quality, then a final: cost and steps
fal lists GPT Image 2.5 at $0.00588 low and $0.05268 high for 1024 squares. Draft four options at low, pick one, then redo it at high. Sume steps and the catch.

Draft at quality: "low" with n: 4, pick the composition you like, then generate the final once at high. fal's price list shows why: a 1024 by 1024 GPT Image 2.5 image costs $0.00588 at low, $0.05268 at high and $0.21072 at max (fal, read 2026-10-01). Four lows cost less than half of one high before margin. Sume bills the endpoint's pricing line, which already includes its margin, so read the live price rather than multiplying fal's.
There is one catch: Sume's image API does not serve a seed, so the high-quality rerun of the same prompt will not reproduce your chosen draft. The workable fix is to pass the draft back as a reference.
What do the quality levels cost?
fal's list rates, before Sume's margin and before input tokens.
| Quality | Output cost |
|---|---|
| low | $0.00588 |
| high | $0.05268 |
| xhigh | $0.09366 |
| max | $0.21072 |
How do I run draft then final?
Step one asks for four low-quality drafts. Step two edits the chosen draft URL at high quality, telling the model to keep the composition. ChatGPT Image 2.5 accepts up to 16 input_references, so the draft is one input among several if you have product shots too.
import os
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
URL = "https://api.sume.com/v1/images"
MODEL = "openai/gpt-image-2.5"
d = requests.post(URL, headers=H, timeout=60, json={
"model": MODEL, "quality": "low", "n": 4,
"prompt": "poster of a red bicycle against a teal wall, bold headline BIKE DAY"})
d.raise_for_status()
for i, x in enumerate(d.json()["data"]):
print(i, x["url"])
pick = int(input("pick: "))
f = requests.post(URL, headers=H, timeout=60, json={
"model": MODEL, "quality": "high", "aspect_ratio": "auto",
"prompt": "Recreate this poster at high fidelity; keep the layout and text",
"input_references": [{"type": "image_url",
"image_url": {"url": d.json()["data"][pick]["url"]}}]})
print(f.status_code, f.json())When does this not pay off?
Low quality is weakest on dense text and fine detail, so a headline that fails at low may be fine at high. If text is the point, draft at medium. If you already know the composition, skip drafts: one high call is cheaper than four lows plus a high. And the final is an edit of the draft, not a fresh render, so it can drift. Compare it against the draft before publishing.
Limits
Omitting quality defaults to high, and auto reserves max, per the docs, so set it explicitly on drafts. n is capped per model; read the range descriptor. Draft URLs must be public HTTPS to be used as references, and Sume-hosted URLs are. Check the live pricing line before budgeting a batch.
Sources
Related posts
More in Use cases
- Keep the same face in a GPT Image 2.5 edit: the prompt block to use
To keep a person recognisable in a GPT Image 2.5 edit, send the photo as a reference and spell out what must not change. The exact wording and a Sume call.
- Halloween product teaser: a 15-second vertical clip from one still
Make a 15-second Halloween teaser for a product with Seedance 2.5 on Sume: first-frame still, 9:16 at 720p, native audio, and how to read the price first.
- Halloween narrator voiceover with spooky music for a video
Make a Halloween story video audio track: a slow narrator from Sume TTS, an instrumental horror bed from the music router, mixed with ducking in Timeline 1.0.
- Higgsfield's Seedance outage on Sept 30: slow vs failed jobs on Sume
Higgsfield said Seedance 2.0 failed more and 2.5 ran slow on Sept 30, now fixed. On Sume, a slow job is queued or processing; rerun only after failed.
Written by Sume