Which AI image model edits a photo best? Reference limits compared

Reference-image limits for edits on GPT Image 2.5, Nano Banana, Grok Imagine Image 2.0 and FLUX, from each vendor's page, and how to call them on Sume.

5 min readSume
All posts

There is no single best editor; the useful difference is how many reference images each model takes and what else it offers for an edit. GPT Image 2.5 gives you up to 16 references plus a mask on Sume, Nano Banana splits references into objects, characters and styles, and Grok Imagine Image 2.0 takes up to 5 source images per request.

The numbers below come from each vendor's own page, read on 2026-10-02, and from the Sume Image API docs. Where a page did not state a limit, the table says so rather than filling it in.

What do the vendor pages say?

OpenAI's guide shows an edit with four input images and states no maximum. Google's guide gives per-model reference mixes. xAI's guide says up to 5 source images in one request, billed for both input and output images. Black Forest Labs' docs home lists editing and reference image combination for FLUX 3 and says FLUX.2 and FLUX.1 Kontext remain supported, without giving a reference limit on that page.

One reading note on the table. The Nano Banana mix is a Google statement about the Gemini models; the reference cap Sume actually applies appears in the catalog, and an earlier post on Nano Banana's 14 references versus Sume's 10 explains one gap. Always trust the descriptor over a vendor page when they differ.

Reference images for edits by model family (read 2026-10-02)
FamilyReference limitOther edit tools
GPT Image 2.5 (Sume docs)Up to 16mask_url, background transparent or opaque
Nano Banana 2 (Google)Up to 10 objects, 4 characters, 3 styleVideo-to-image, search grounding
Nano Banana Pro (Google)Up to 6 objects, 5 characters, 3 styleSearch grounding
Nano Banana 2 Lite (Google)Up to 14 object images1K only
Grok Imagine Image 2.0 (xAI)Up to 5 source imagesPublic URL or base64 input
FLUX family (BFL)Not stated on the page readOutpainting, erase, deblur, virtual try-on

Which one for which edit?

Match the edit to the feature, not the brand.

  • Fix one region of a photo: GPT Image 2.5 with mask_url.
  • Keep a character or product consistent across scenes: a Nano Banana row, where characters and objects are separate budgets.
  • Combine a few subjects cheaply: Grok Imagine Image 2.0, remembering each input image is billed.
  • Expand a canvas or remove an object: check the FLUX editing tools in the catalog before assuming a Sume row exposes them.

How do you call them the same way on Sume?

One endpoint serves all of them, and the model id is the only change. Read each model's input_references descriptor first: a request that sets a parameter the model does not list returns 400 unsupported_parameter, and models whose descriptor is min 0, max 0 are text-to-image only. Sume's limit can be lower than the vendor's, so the catalog is the final answer.

Costs follow the same split. GPT Image 2.5 on Sume is charged from token rates, so every reference image adds input tokens. Grok's input images are billed per the xAI guide. Check the pricing lines before sending sixteen references when two would do.

import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
r = requests.get("https://api.sume.com/v1/images/models", headers=H, timeout=30)
for m in r.json()["data"]:
    refs = m["supported_parameters"].get("input_references")
    if refs and refs.get("max", 0) > 0:
        print(f'{m["id"]:55} refs {refs["min"]}-{refs["max"]}'
              f' mask={"mask_url" in m["supported_parameters"]}')

What does the table leave out?

Quality. None of the pages say which model keeps a face or a logo best, and no number here ranks them. Run your own hard case through two or three rows and compare. Our compare script post shows how to run one prompt on several models.

Sources

Related posts

More in Models

All Models posts

Written by Sume