Nano Banana reference limits by model: objects, characters, style

Google splits Nano Banana references by model: Nano Banana 2 takes 10 objects, 4 characters, 3 style images; Pro takes 6, 5, 3. Sume uses one list.

5 min readSume
All posts

Google's current image generation page lists reference limits per model: Nano Banana 2 (gemini-3.1-flash-image) takes 10 high-fidelity objects, 4 characters and 3 style references; Nano Banana Pro (gemini-3-pro-image) takes 6 objects, 5 characters and 3 style references; and Nano Banana 2 Lite takes 14 objects with no character or style slots listed. Sume's image API takes a single input_references list, which the catalog caps per model, and does not label slots as object, character or style.

The limits are from Google's Nano Banana image generation page, read 2026-10-02. Sume's side is in the Images API docs. If you read an older figure of 14 objects and 5 characters for Nano Banana 2, the page I fetched today splits that by model as below.

What does Google list for each model?

The table copies the page's reference table. "N/A" means the page lists nothing for that slot, not that the model refuses the image.

Maximum reference images by Nano Banana model, read 2026-10-02
ModelAPI idHigh-fidelity objectsCharactersStyle references
Nano Banana 2 Litegemini-3.1-flash-lite-image14N/AN/A
Nano Banana 2gemini-3.1-flash-image1043
Nano Banana Progemini-3-pro-image653

How does this differ from what Sume accepts?

Sume describes each image model with typed capability descriptors. For models that take reference images, input_references is a range descriptor with min and max, and the docs' catalog example shows { "min": 0, "max": 10 }. A request that exceeds the descriptor, or sets a parameter the model does not list, is rejected with 400 unsupported_parameter instead of being silently trimmed.

There is no per-slot split. If you send four character photos, two product shots and a style board, that is seven entries in one list, and Sume forwards them without telling the model which is which. Describe the roles in the prompt instead, for example "the person in image 1, the jacket in image 2", and keep the total under the descriptor.

How should you split a large reference set?

Rank by what must stay identical. Characters and the key product usually come first, style references last. With Nano Banana 2 on Google's side that means at most 4 character images and 3 style images, so a set of 6 people does not fit in one request there. On Sume the same set fits a 10-entry list, but Google's character limit still describes what the model was built to hold, so consistency may drop beyond it. Run a shot with the first four people, then a second shot that adds the rest, and compare.

Lite is the odd one: 14 object references but no character or style slots and 1K only. Sume's docs do not list a Lite id, so treat it as Google-only.

What does a Sume request with several references look like?

Reference URLs must be public HTTPS; localhost and private-network URLs are rejected before submission. Use aspect_ratio: "auto" on edits to keep the reference shape.

curl -X POST https://api.sume.com/v1/images \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nano-banana-2",
    "prompt": "The woman from image 1 wearing the jacket from image 2, city street, 35mm look",
    "aspect_ratio": "auto",
    "input_references": [
      {"type": "image_url", "image_url": {"url": "https://example.com/person.jpg"}},
      {"type": "image_url", "image_url": {"url": "https://example.com/jacket.jpg"}}
    ]
  }'

How do output limits fit in?

Reference counts are one half of the request. Google's page also lists the output side: Nano Banana 2 offers 512px, 1K, 2K and 4K, Pro offers 1K, 2K and 4K, and Lite is 1K only, with ten aspect ratios from 1:1 to 21:9. Sume's docs list a normalized resolution tier (512, 1K, 2K, 4K) and a wider ratio set, but a model accepts only what its descriptors list, so read them first.

Google also notes that thinking cannot be disabled in the API and has two levels, minimal (the default) and high, on the Flash and Flash Lite models. Sume's request has no thinking field in the parameter table I read, so that choice stays on Google's side.

What should you check before relying on a limit?

Read GET /v1/images/models for the live descriptor before pinning a number into your code, and read Google's page again when a model id changes. Both vendors move limits between releases, and a number copied from a blog post, including this one, is only as current as its read date.

Sources

Related posts

More in Models

All Models posts

Written by Sume