Nano Banana reference limits by model: objects, characters, style
Google splits Nano Banana references by model: Nano Banana 2 takes 10 objects, 4 characters, 3 style images; Pro takes 6, 5, 3. Sume uses one list.

Google's current image generation page lists reference limits per model: Nano Banana 2 (gemini-3.1-flash-image) takes 10 high-fidelity objects, 4 characters and 3 style references; Nano Banana Pro (gemini-3-pro-image) takes 6 objects, 5 characters and 3 style references; and Nano Banana 2 Lite takes 14 objects with no character or style slots listed. Sume's image API takes a single input_references list, which the catalog caps per model, and does not label slots as object, character or style.
The limits are from Google's Nano Banana image generation page, read 2026-10-02. Sume's side is in the Images API docs. If you read an older figure of 14 objects and 5 characters for Nano Banana 2, the page I fetched today splits that by model as below.
What does Google list for each model?
The table copies the page's reference table. "N/A" means the page lists nothing for that slot, not that the model refuses the image.
| Model | API id | High-fidelity objects | Characters | Style references |
|---|---|---|---|---|
| Nano Banana 2 Lite | gemini-3.1-flash-lite-image | 14 | N/A | N/A |
| Nano Banana 2 | gemini-3.1-flash-image | 10 | 4 | 3 |
| Nano Banana Pro | gemini-3-pro-image | 6 | 5 | 3 |
How does this differ from what Sume accepts?
Sume describes each image model with typed capability descriptors. For models that take reference images, input_references is a range descriptor with min and max, and the docs' catalog example shows { "min": 0, "max": 10 }. A request that exceeds the descriptor, or sets a parameter the model does not list, is rejected with 400 unsupported_parameter instead of being silently trimmed.
There is no per-slot split. If you send four character photos, two product shots and a style board, that is seven entries in one list, and Sume forwards them without telling the model which is which. Describe the roles in the prompt instead, for example "the person in image 1, the jacket in image 2", and keep the total under the descriptor.
How should you split a large reference set?
Rank by what must stay identical. Characters and the key product usually come first, style references last. With Nano Banana 2 on Google's side that means at most 4 character images and 3 style images, so a set of 6 people does not fit in one request there. On Sume the same set fits a 10-entry list, but Google's character limit still describes what the model was built to hold, so consistency may drop beyond it. Run a shot with the first four people, then a second shot that adds the rest, and compare.
Lite is the odd one: 14 object references but no character or style slots and 1K only. Sume's docs do not list a Lite id, so treat it as Google-only.
What does a Sume request with several references look like?
Reference URLs must be public HTTPS; localhost and private-network URLs are rejected before submission. Use aspect_ratio: "auto" on edits to keep the reference shape.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2",
"prompt": "The woman from image 1 wearing the jacket from image 2, city street, 35mm look",
"aspect_ratio": "auto",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/person.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/jacket.jpg"}}
]
}'How do output limits fit in?
Reference counts are one half of the request. Google's page also lists the output side: Nano Banana 2 offers 512px, 1K, 2K and 4K, Pro offers 1K, 2K and 4K, and Lite is 1K only, with ten aspect ratios from 1:1 to 21:9. Sume's docs list a normalized resolution tier (512, 1K, 2K, 4K) and a wider ratio set, but a model accepts only what its descriptors list, so read them first.
Google also notes that thinking cannot be disabled in the API and has two levels, minimal (the default) and high, on the Flash and Flash Lite models. Sume's request has no thinking field in the parameter table I read, so that choice stays on Google's side.
What should you check before relying on a limit?
Read GET /v1/images/models for the live descriptor before pinning a number into your code, and read Google's page again when a model id changes. Both vendors move limits between releases, and a number copied from a blog post, including this one, is only as current as its read date.
Sources
Related posts
More in Models
- Nano Banana text in images: write the copy first, then render
Google's tip for text in Gemini images: settle the wording first, then ask for the image. How to do that with one Sume request and a short checklist.
- OpenAI media model shutdown calendar 2026: DALL-E, Sora, GPT Image
Three OpenAI media retirements in 2026: DALL-E on May 12, Sora and the Videos API on September 24, GPT Image 1.5 and mini on December 1. Dates and replacements.
- Qwen-Image 2.0 Pro is Alibaba's pick: which Qwen ids does Sume list?
Alibaba recommends qwen-image-2.0-pro. Sume's image catalog lists qwen/qwen-image and qwen/qwen-image-max in the repo; check the live catalog for more.
- Qwen-Image negative_prompt 500 characters vs a Sume image request
Alibaba's Qwen-Image API takes a negative_prompt up to 500 characters. Sume's image request table has no such field, so rewrite exclusions into the prompt.
Written by Sume