12 reference photos in one edit: only GPT Image 2.5 takes them
Nano Banana 2.1 takes 10 references on Sume, Ideogram 4.5 takes 5, GPT Image 2.5 takes 16. For a 12-photo edit use GPT Image 2.5, or split the job in two.

If one edit needs 12 reference photos, openai/gpt-image-2.5 is the Sume row that takes them: its input_references range is 0 to 16. Nano Banana 2.1 stops at 10, and Ideogram 4.5 at 5. Tencent's Hy Image 3.5 Preview lists up to 20 on OpenRouter, but Sume does not list it.
Limits side by side
Sume's shared reference ceiling is 10, with per-model exceptions of 16 for both ChatGPT Image 2.5 ids and 5 for Ideogram 4.5. Imagen 4 takes none.
| Model | Max references | Where |
|---|---|---|
| openai/gpt-image-2.5 | 16 | Sume catalog |
| google/nano-banana-2.1 | 10 | Sume catalog |
| ideogram/ideogram-v4.5 | 5 | Sume catalog |
| google/imagen-4-fast, imagen-4-ultra | 0 | Sume catalog |
| Hy Image 3.5 Preview | 20 | OpenRouter page; not on Sume |
The price catch
GPT Image 2.5 is billed on tokens, and each reference adds input tokens. The docs give $8 per million input image tokens at the provider and say input counts are estimates. With no size or quality set, a text-to-image call quotes $0.2225 on Sume and a one-reference edit $0.28175. Set image_size and quality to keep a 12-photo edit predictable.
Splitting instead
If 10 references are enough once you pick the best ones, use Nano Banana 2.1 at $0.10 for 1K. Rank your photos, keep the 10 that show the product label, color and shape, and drop duplicates. Fewer, cleaner references usually beat a pile of near-identical ones, but test on your product before you trust that.
Choosing the references
Sort the 12 photos by what each shows: two for the overall shape, two for the label, two for color accuracy, two for texture and four for context. If you must cut to 10, drop one from context and one from texture first. The label and the shape matter most.
If you go with GPT Image 2.5, also pass mask_url when you want to limit the edit to part of the image, which that row lists. Nano Banana 2.1 does not list a mask field, and a request with one returns a 400.
Compare on one test product before you commit a catalog: run the 12-reference GPT edit and a 10-reference Nano Banana 2.1 edit on the same item and look at the label.
What the numbers rely on
Reference limits come from Sume's catalog code: a shared ceiling of 10, with 16 for both GPT Image 2.5 ids and 5 for Ideogram 4.5. The Hy Image figure of 20 is from OpenRouter's model page. Treat 20 as the OpenRouter listing, not as a Sume limit.
The code can change. Read supported_parameters.input_references from the live catalog.
If your photos are public on a product page, you can pass those URLs directly. If they sit behind a login, host copies on a public HTTPS path first. Sume rejects localhost, private-network and non-HTTPS URLs before submission, so a failed submit on a good-looking URL is usually a private host. Test one URL with a single edit call before you queue a whole set.
Sources
Related posts
More in Comparisons
- 12 s at 720p: Seedance 2.5 $6.94, 2.0 $4.54, Fast $3.63, Mini $2.27
Four Seedance rows priced for the same 12-second 720p clip on Sume, with the list rate per 1000 tokens and when to pick each.
- 12 s clip: sume/auto caps at 10 s; pin Wan $1.50 or Seedance 2.5 $6.94
sume/auto takes 3 to 10 s, so a 12 s clip needs a pinned model. 12 s at 720p on Sume: H3 $0.90, Wan $1.50, Kling $1.68, Mini $2.27, Seedance 2.5 $6.94.
- 15-second clip with sound: Kling 3 $3.15, Wan 3.0 $1.88, Mini $2.84
A 15 s 720p clip with audio on Sume: Kling 3 with sound is $3.15, Wan 3.0 $1.88, Seedance 2 Mini $2.84 (9:16), MiniMax H3 $1.13. Arithmetic and limits.
- 20-second talking clip: Fabric 720p $3.75 vs Avatar Standard $3.68
For 20 seconds, VEED Fabric 720p on Sume bills $3.75 (20 x $0.1875) and Avatar Video Standard $3.68 (20 x $0.184). The inputs differ: audio + still vs a script.
Written by Sume