Image reference limits: GPT Image 2.5 takes 16, Ideogram 4.5 takes 5
GPT Image 2.5 accepts up to 16 input references on Sume; Ideogram 4.5 edits the first image and uses up to 4 more. How to order them and when to prefer which.

On Sume, openai/gpt-image-2.5 and openai/gpt-image-2.5-sunburst accept up to 16 input_references, while ideogram/ideogram-v4.5 takes 5 in total: with references it edits the first image and uses up to 4 more as references. If your job is combining many assets, such as a model plus several garments, 16 gives room. If it is editing one picture with a few style cues, 5 is plenty.
What the docs say
The Image API page states both limits: Flare and Sunburst support text-to-image, up to 16 image references and an optional mask_url; Ideogram 4.5 generates from text without references, and with them edits the first image. An edit without aspect_ratio keeps the source image's shape. Other families list their own input_references range, and a model with {"min": 0, "max": 0} is text-to-image only and rejects references.
OpenAI's image generation guide describes the same capability in its own words: one or more images as references for new creations, plus mask-based editing.
| Model id | Max references | How the first image is used |
|---|---|---|
| openai/gpt-image-2.5 | 16 | Reference; edit and compose with the others |
| openai/gpt-image-2.5-sunburst | 16 | Same limits as Flare |
| ideogram/ideogram-v4.5 | 5 total | First image is edited; up to 4 more are references |
| bytedance-seed/seedream-4.5 | 10 (docs example) | Per the catalog descriptor shown in the docs |
Order matters
The prompt should name references by position: image 1 is the person, images 2 to 4 are garments. On Ideogram the first position is special because it is the image being edited, so put the picture you want changed there and say what the other four are for. Keep the list short: send only the references the prompt names.
{
"model": "ideogram/ideogram-v4.5",
"prompt": "Edit image 1: swap the mug for the teapot in image 2. Keep the table and light.",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/table.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/teapot.jpg"}}
],
"quality": "medium"
}Rules for every reference
Reference URLs must be public HTTPS. Localhost, private-network and non-HTTPS URLs are rejected before submission, as the Image API docs say, and the Media inputs page covers URL rules for uploads. Ask a text-only model for a reference and you get a rejection, not a quiet ignore.
Picking between them
- Many assets into one scene (outfits, kits, collages): GPT Image 2.5, up to 16.
- One photo to retouch with a couple of style cues: Ideogram 4.5 keeps the source shape by default.
- Local fix to one region: GPT Image 2.5 with
mask_url. - Unsure: send the same two-image edit to both at
n: 1and compare the results andusage.cost.
Sources
Related posts
More in Models
- Index-Translate 2B, 9B or 35B-A3B for subtitle translation
Index-Translate ships 2B, 9B and 35B-A3B text models. What the vendor pages say about each size, and how to pick one for subtitle cues you burn with Sume.
- Index-Translate, Echo, Homura: which part do you need?
Bilibili's Index-Translate release has text, speech, dubbing and length parts. Which fit subtitles, and which Sume step follows.
- Inworld TTS 2 style steering and the Sume voiceover path
What is reported about Inworld TTS 2 style steering and 100+ languages, and how to produce a voiceover with Sume's tts_create tool and join takes.
- Nano Banana 3: does it exist? Current ids
Google autocomplete suggests a Nano Banana 3, but suggestions are not releases. What autocomplete lists and which Nano Banana id Sume accepts.
Written by Sume