Draft with Qwen Image, finish with ChatGPT Image 2.5: 50 heroes cost
Four Qwen Image drafts at $0.025 then one ChatGPT Image 2.5 finish per hero costs $8.29375 for 50 heroes, against $13.175 for four GPT attempts each.

A two-model pipeline is cheaper than rerolling the expensive one: generate four Qwen Image drafts per hero at $0.025 each, pick one, and send it as a reference to ChatGPT Image 2.5 for the finish. For 50 hero images that is $8.29375 against $13.175 if you take four GPT attempts per hero.
The saving comes from the draft stage, where you buy composition choices at a quarter of the price. It only holds if the finishing model keeps the draft's layout, which you should test on five heroes before running 50.
The cost model
Both prices are the billed rates from the Sume catalog (read 2026-10-03); the ChatGPT Image 2.5 rate is the high-quality 1024 figure, and a larger size or xhigh quality raises it.
| Plan | Calls | Images billed | Total |
|---|---|---|---|
| Qwen drafts (n=4) then one GPT 2.5 finish | 50 + 50 | 200 Qwen + 50 GPT | $8.29375 |
| GPT 2.5 only, four attempts per hero | 50 x 4 | 200 GPT | $13.175 |
| GPT 2.5 only, one attempt per hero | 50 | 50 GPT | $3.29375 |
Why the third row is not the answer
One attempt per hero is the cheapest line, but it leaves you with whatever the first sample was. The draft stage is a cheap selection step: four compositions, one chosen by a human or a scoring rule, then a paid finish on the winner. The finish call uses the draft as an input_references item, so the composition carries over while the finishing model redraws detail.
ChatGPT Image 2.5 accepts up to 16 references and Qwen Image 10, so you can also pass a style guide next to the draft.
The two calls
Both calls use POST /v1/images. The helper raises on a 202 so a slow generation is not mistaken for an image; switch to mode: "async" and the job endpoints for 4K or high-quality finishes.
import os
import requests
URL = "https://api.sume.com/v1/images"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def gen(payload):
r = requests.post(URL, headers=H, json=payload, timeout=60)
r.raise_for_status()
if r.status_code != 200:
raise RuntimeError("202: poll the job instead")
return r.json()["data"]
def hero(prompt):
drafts = gen({"model": "qwen/qwen-image", "prompt": prompt, "n": 4})
pick = drafts[0]["url"] # replace with your own choice
ref = {"type": "image_url", "image_url": {"url": pick}}
return gen({
"model": "openai/gpt-image-2.5",
"prompt": "Finish image 1 as a polished hero: " + prompt,
"input_references": [ref],
})[0]["url"]
if __name__ == "__main__":
print(hero("a ceramic mug on a walnut desk, soft window light"))Sources
Related posts
More in Use cases
- Dub a video, keep the music: Sume has no stem splitter, so do this
Sume can detach a video's audio and re-voice a script, but its docs list no vocal and music separation. A clean workaround with a new music bed and ducking.
- Dub a video where two languages are spoken: keep the original lines
Sume's speech-to-text takes one language hint per call. For a mixed-language video, transcribe, translate only some lines, and splice the rest back in.
- DV360 video creative specs: H.264, 20 Mbps, -24 LKFS audio
Display & Video 360 asks for H.264 at 20 Mbps or more, 23.98 or 29.97 fps, 48 kHz audio at -24 LKFS. Which of those Sume's trim and timeline can and cannot set.
- eBay PictureURL: first URL is the Gallery image, 3,975-char cap
eBay allows up to 24 picture URLs, uses the first as the Gallery image, and caps all PictureURL values at 3,975 characters. How to order and size a Sume batch.
Written by Sume