Image-to-image in Python: three references on GPT Image 2.5
A 22-line Python script that sends three reference images to POST /v1/images with GPT Image 2.5 and prints the URL and cost. References limits and pitfalls.

To do image-to-image on Sume with ChatGPT Image 2.5, send a prompt and an input_references array of image objects to POST /v1/images; the row accepts up to 16 references. The script below sends three product photos and asks for a gift-basket arrangement. It prints the first image URL and the usage.cost figure.
The script
import json, os, urllib.request
refs = [
"https://example.com/soap.jpg",
"https://example.com/candle.jpg",
"https://example.com/tea.jpg",
]
body = {
"model": "openai/gpt-image-2.5",
"prompt": "a gift basket holding the three items shown, on a wooden table",
"quality": "medium",
"aspect_ratio": "4:3",
"input_references": [
{"type": "image_url", "image_url": {"url": u}} for u in refs
],
}
req = urllib.request.Request(
"https://api.sume.com/v1/images", json.dumps(body).encode(),
{"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=60) as r:
out = json.load(r)
print(r.status, out["data"][0]["url"], out["usage"]["cost"])Rules that make it work
Reference URLs must be public HTTPS; Sume rejects localhost, private-network and non-HTTPS URLs before submission. Each entry is an object with type: "image_url" and an image_url.url, not a bare string. If the response status is 202 rather than 200, the body is a job envelope without data[0].url, so the print line would fail; see the status-code branch in the Image API docs.
| Field | Value in the script | Why |
|---|---|---|
| input_references | 3 image_url objects | Image-to-image; up to 16 on this row |
| quality | medium | 2 cents text price; edits add input tokens |
| aspect_ratio | 4:3 | Listed by the row; use auto to match one reference |
| model | openai/gpt-image-2.5 | Flare; Sunburst is the other id |
What it costs
Fal's token rates, which the Sume docs say apply to both 2.5 ids, are $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens. Input token counts are estimates, and the total is rounded up to $0.0001. So three references add input tokens on top of the output charge, and the exact amount appears in usage.cost after the call.
Do not budget a reference-heavy edit at the text-only price. Run one call, read the cost, then multiply.
Pitfalls
A few details decide whether an edit like this works first time.
- Flare and Sunburst share the same price and limits; OpenAI positions Sunburst for editing precision, so test it if edits drift.
- More references is not always better; each adds input cost and can dilute the prompt.
- State what each reference is for in the prompt: 'the soap from image 1'.
- Keep references near the target aspect ratio to avoid odd crops.
Checking the response
The script prints the HTTP status, the first URL and the cost. A 200 means you have the image; URLs are Sume-hosted and signed, so download them promptly if you need a permanent copy. If you send a mask, add mask_url as another top-level field in the body; if you want the output to match the first reference's shape, replace the ratio with auto. Put each reference's role in the prompt, because the model sees them as a set, not as ordered instructions.
Choosing Flare or Sunburst
Both ids take the same request and cost the same on Sume. OpenAI's guide describes Sunburst as the one for workflows where editing precision matters most and Flare as the fast, high-quality everyday choice. For a multi-reference composition, run the same three references through each at medium quality and compare; the extra cost is one more call. Auto routing uses Flare, so pin the Sunburst id explicitly if it wins your test.
Keep the three reference URLs reachable until the job is complete. Sume fetches public HTTPS URLs at submission, and a URL that returns an error will fail the call before any image is made.
Sources
Related posts
More in Developers
- image_url or reference_image_urls: the Omni field for a photo
On Gemini Omni Flash 1.1, one image in image_url is a start frame; one image in reference_image_urls with no frame is reference-to-video. The price is the same.
- Imagen 4 with a reference image: Sume rejects it, use Nano Banana 2.1
Sume lists Imagen 4 Fast and Ultra as text-to-image only: input_references is 0 to 0. To edit a photo, use Nano Banana 2.1 at $0.10 per 1K image.
- Why input_references fails on Imagen, Recraft, Soul and Qwen Max
Five Sume image rows list zero input_references: imagen-4 fast and ultra, recraft-v4, Soul and qwen-image-max. Check the catalog in Python first.
- Is Higgsfield Genjutsu in your Sume catalog? Check, then price it
higgsfield-genjutsu appears in the catalog only when its provider is configured. List your models, then price a 4 to 30 second source at 480p or 720p.
Written by Sume