Which AI image model edits a photo best? Reference limits compared
Reference-image limits for edits on GPT Image 2.5, Nano Banana, Grok Imagine Image 2.0 and FLUX, from each vendor's page, and how to call them on Sume.

There is no single best editor; the useful difference is how many reference images each model takes and what else it offers for an edit. GPT Image 2.5 gives you up to 16 references plus a mask on Sume, Nano Banana splits references into objects, characters and styles, and Grok Imagine Image 2.0 takes up to 5 source images per request.
The numbers below come from each vendor's own page, read on 2026-10-02, and from the Sume Image API docs. Where a page did not state a limit, the table says so rather than filling it in.
What do the vendor pages say?
OpenAI's guide shows an edit with four input images and states no maximum. Google's guide gives per-model reference mixes. xAI's guide says up to 5 source images in one request, billed for both input and output images. Black Forest Labs' docs home lists editing and reference image combination for FLUX 3 and says FLUX.2 and FLUX.1 Kontext remain supported, without giving a reference limit on that page.
One reading note on the table. The Nano Banana mix is a Google statement about the Gemini models; the reference cap Sume actually applies appears in the catalog, and an earlier post on Nano Banana's 14 references versus Sume's 10 explains one gap. Always trust the descriptor over a vendor page when they differ.
| Family | Reference limit | Other edit tools |
|---|---|---|
| GPT Image 2.5 (Sume docs) | Up to 16 | mask_url, background transparent or opaque |
| Nano Banana 2 (Google) | Up to 10 objects, 4 characters, 3 style | Video-to-image, search grounding |
| Nano Banana Pro (Google) | Up to 6 objects, 5 characters, 3 style | Search grounding |
| Nano Banana 2 Lite (Google) | Up to 14 object images | 1K only |
| Grok Imagine Image 2.0 (xAI) | Up to 5 source images | Public URL or base64 input |
| FLUX family (BFL) | Not stated on the page read | Outpainting, erase, deblur, virtual try-on |
Which one for which edit?
Match the edit to the feature, not the brand.
- Fix one region of a photo: GPT Image 2.5 with
mask_url. - Keep a character or product consistent across scenes: a Nano Banana row, where characters and objects are separate budgets.
- Combine a few subjects cheaply: Grok Imagine Image 2.0, remembering each input image is billed.
- Expand a canvas or remove an object: check the FLUX editing tools in the catalog before assuming a Sume row exposes them.
How do you call them the same way on Sume?
One endpoint serves all of them, and the model id is the only change. Read each model's input_references descriptor first: a request that sets a parameter the model does not list returns 400 unsupported_parameter, and models whose descriptor is min 0, max 0 are text-to-image only. Sume's limit can be lower than the vendor's, so the catalog is the final answer.
Costs follow the same split. GPT Image 2.5 on Sume is charged from token rates, so every reference image adds input tokens. Grok's input images are billed per the xAI guide. Check the pricing lines before sending sixteen references when two would do.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
r = requests.get("https://api.sume.com/v1/images/models", headers=H, timeout=30)
for m in r.json()["data"]:
refs = m["supported_parameters"].get("input_references")
if refs and refs.get("max", 0) > 0:
print(f'{m["id"]:55} refs {refs["min"]}-{refs["max"]}'
f' mask={"mask_url" in m["supported_parameters"]}')What does the table leave out?
Quality. None of the pages say which model keeps a face or a logo best, and no number here ranks them. Run your own hard case through two or three rows and compare. Our compare script post shows how to run one prompt on several models.
Sources
Related posts
More in Models
- Zalando designer minimum 1800x2600: request 1808x2608 on Sume
Zalando designer brands need at least 1800x2600 px in 1:1.44 JPEG. GPT custom sizes need multiples of 16, so ask Sume for 1808x2608, which clears both edges.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
Written by Sume