Nano Banana 2.1 references: 10 objects, 4 characters vs Pro 6, 5, 3
Google splits Nano Banana 2.1 references into 10 objects and 4 characters, and Pro into 6 objects, 5 characters and 3 style images. How Sume's flat list maps.

Google's image guide says Nano Banana models take up to 14 reference images, but the 14 is split by role. Nano Banana 2.1 allows up to 10 object images and up to 4 character images. Pro allows up to 6 object images, 5 character images and 3 style references. Sume's image API takes one flat input_references list, so it has no field that tells the model which role an image plays.
Google's limits by role
The table is from Google's reference-image section. "N/A" means the guide lists no allowance for that role on that model.
| Role | Nano Banana 2.1 | Nano Banana Pro |
|---|---|---|
| High-fidelity objects | Up to 10 | Up to 6 |
| Characters for consistency | Up to 4 | Up to 5 |
| Style references | N/A | Up to 3 |
| Total mixed | Up to 14 | Up to 14 |
What changes for you
- A 2.1 job with 10 products and 4 people uses the full 14 on Google's side.
- Pro takes fewer objects but accepts a style image. For a brand look from a reference, Pro is the only one of the two that lists it.
- Each role has its own cap, so five characters is over the limit on 2.1 even if the total is under 14.
On Sume
In the repo's catalog code, every edit-capable image model gets 10 input_references unless it has its own entry, and Nano Banana 2.1 and Pro have none. The Image API docs show references as a list of image URLs and say a request outside the descriptor is rejected with 400 unsupported_parameter. Describe each image's role in the prompt, and read supported_parameters.input_references for your key before you plan a 14-image job. I did not find a Sume field for role-tagged references.
Sources
Related posts
More in Models
- Nano Banana 2 is retired on Sume: same price on Nano Banana 2.1?
Sume retired google/nano-banana-2 and runs old requests as Nano Banana 2.1 at the same price card: $0.10 at 1K billed. Tiers, ids and what the job stores.
- Native 4K or 2K image edits: pixel ceilings by model
How many pixels each model reaches: Flux 3 Image near 16.8 MP, GPT Image 2.5 capped at 8,294,400 on Sume, Ideogram 4.5 at 2K. What Sume serves.
- Gemini Omni 'no dialogue' prompt: steer the audio on Sume
Omni clips always carry sound. Google suggests 'no dialogue' and 'no extra sound effects' in the prompt. How to use that on Sume and what a test costs.
- Gemini Omni sound effects prompt: name each sound in 8 seconds
How to prompt footsteps, a door and breaking glass in a Gemini Omni clip: Google says describe audio explicitly. Sume request, timing and cost per take.
Written by Sume