GPT Image 2.5 edit with several references: number each image
With up to 16 references in a GPT Image 2.5 edit, name each one by number and role in the prompt. A Sume request with three references, and what to check.

When a GPT Image 2.5 edit carries more than one reference image, number them in the prompt and give each a single job: "Image 1 is the person, Image 2 is the product, Image 3 is the room." Both OpenAI's prompting guide and fal's GPT Image 2.5 guide say to do this, and on Sume you send the images in input_references of POST /v1/images with openai/gpt-image-2.5 (Flare) or openai/gpt-image-2.5-sunburst.
The sources are OpenAI's Image prompting guide and fal's How To Use GPT Image 2.5, both read on 2026-10-02, plus Sume's Image API page. Sume's docs say both GPT Image 2.5 ids accept up to 16 image references. They do not say how the array order maps to the words "Image 1", so the habit below is to keep the order of your array identical to your numbering and to check the first result.
Why number the references at all?
Without numbers the model has to guess which picture supplies the face, which supplies the product and which only sets the mood. OpenAI's guide puts it as identifying each input by number and purpose. fal's guide gives the same advice with an example form, "Image 1 is...", and says to assign each image a specific task.
A vague instruction leaves the scope of the change to the model. A numbered instruction settles it before the render, which is cheaper than a retry. Each completed generation is billed in full on Sume, so a second try costs a second image.
What does a numbered request look like on Sume?
This request puts a portrait, a product shot and a room photo into one edit. The prompt names each by number, says what each contributes, and lists what must stay as it is. The three URLs must be public HTTPS URLs; localhost, private-network and non-HTTPS URLs are rejected before submission.
curl -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Image 1 is the person: keep her face and pose exactly. Image 2 is the product: keep its shape, colour and label text exactly. Image 3 is the room. Show the person from Image 1 holding the product from Image 2 in the room from Image 3, natural window light.",
"quality": "high",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/person.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/product.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/room.jpg"}}
]
}'What should each reference be assigned?
One role per image keeps the prompt readable. The table is a starting pattern, not a Sume rule.
| Slot | Role | Constraint line to add |
|---|---|---|
| Image 1 | Subject (a person or a character) | Keep the face and pose exactly |
| Image 2 | Product or object | Keep shape, colour and label text exactly |
| Image 3 | Scene or background | Use as the setting only, do not copy its people |
| Prompt tail | Style and light | One sentence, after the three roles |
What can go wrong with many references?
More references do not guarantee that all of them are honoured. fal's guide notes that repeated edits can still shift details you meant to keep, and the same caution applies to a crowded first pass. Give the prompt fewer jobs, put the hardest constraint first, and inspect the result against each source image.
If you also need to confine the change to one region, a mask_url is available on GPT Image 2.5 edits. Sume's docs list it as an optional public HTTPS URL; see how a mask behaves with several references.
Sume does not publish a fidelity score for any reference, and usage.cost on the response is the only per-request number it returns beyond the image URLs. Treat the first render as a test of your numbering.
- Keep the array order identical to the numbers in the prompt.
- Use one role per reference.
- State what must not change for each image that carries a subject.
- Check the render against every source image, not only the first.
What is a quick test of the numbering?
Run the same request twice, once with the numbered prompt and once with the references described only in prose, and compare. If the numbered run puts the right subject in the right place and the prose run does not, keep the numbers. If both are right, you saved nothing by numbering, but the habit costs one sentence.
Flare (openai/gpt-image-2.5) is the faster of the two ids and Sunburst (openai/gpt-image-2.5-sunburst) trades time for detail, per fal's guide; Sume's Auto routing continues to use Flare. Test your numbering on Flare first, and move to Sunburst only for the final render, so a numbering mistake costs the cheaper image. The two ids use the same token rates in Sume's docs.
Sources
- Image API
- Image prompting (OpenAI, read 2026-10-02)
- [How To Use GPT Image 2.5: Prompts & Workflows [2026] (fal, read 2026-10-02)](https://fal.ai/learn/tools/how-to-use-gpt-image-2-5)
Related posts
More in Developers
- Graph API v26.0: does it change Reels publishing? Not by its notes
Graph API v26.0 shipped July 29, 2026. Its changelog lists no Reels or video changes, but removes Explore placement and five Page fields. Dates inside.
- Grok Imagine extension: duration is added seconds; Sume's chain
xAI's /v1/videos/extensions duration counts only the new seconds, so 10s plus 5 returns 15s. Sume has no extend call here: chain clips from the last frame.
- MiniMax H3 Max lip sync audio_url: Sume media host only
Sume's MiniMax H3 Max Lip Sync takes a public HTTPS audio_url on the Sume media host, 5 to 14.8 seconds, max 10 MB. Other hosts are rejected.
- MiniMax H3 Max lip sync: 1080p max, no 2K, speed_tier ignored
Sume's MiniMax H3 Max Lip Sync offers 480p, 768p (default) and 1080p; 2K is not offered, and speed_tier is accepted but ignored.
Written by Sume