Nano Banana 2: five characters, 14 objects vs Sume reference slots
Google says Nano Banana 2 keeps up to five characters and 14 objects consistent. Sume caps reference images per model; see what you can send and where it stops.

Google says Nano Banana 2 can 'maintain character resemblance of up to five characters and the fidelity of up to 14 objects in a single workflow' (Google, read 2026-10-01). On Sume you do not set a character count. You pass reference images in input_references, and the ceiling is the per-model range your catalog row reports, which is 10 by default and 16 on GPT Image 2.5.
So the question to ask is how many images you can attach, not how many characters exist. Read the descriptor from GET /v1/images/models before building a storyboard tool around a number from a launch post.
How do the two compare?
Vendor claim and Sume limit, side by side.
| Item | Sume | |
|---|---|---|
| Characters kept consistent | Up to 5 (workflow) | No character setting |
| Objects kept consistent | Up to 14 (workflow) | No object setting |
| Reference images per request | Not stated in the post | 10 default, 16 on GPT Image 2.5, 0 on text-only |
| Where to read it | Google blog | input_references range in the model list |
How do I read the limit?
A short script prints the cap for the Banana rows.
import os
import requests
r = requests.get(
"https://api.sume.com/v1/images/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
if "banana" in m["id"]:
params = m.get("supported_parameters", {})
print(m["id"], params.get("input_references"))What should I do with a cast of characters?
Put each character on its own clean reference, describe who is who in the prompt ('the person from image 1'), and keep the number of references at or below the printed cap. For a longer story, generate a few frames per call and reuse the best frame as a reference in the next one. See Nano Banana Pro vs Nano Banana 2 for the other Banana row.
A practical storyboard recipe
Start with one reference per character on a plain background, then write the prompt so each person is named by image position. Generate one frame, review it, and only then add props as extra references. If a face drifts, remove other references from that call rather than adding more, because fewer references leaves less to confuse.
Keep a small table of which reference URL is which character so your script can rebuild the list for every frame. This makes the cap visible: if your cast plus props is above the printed limit, split the scene into two calls and composite, or drop the least important prop.
Limits
The five-and-14 figures are Google's description of its own workflow, and I did not reproduce them through Sume. The model list's field names may differ slightly from this script, so print the whole row once if the lookup returns None. References must be public HTTPS URLs.
Sources
Related posts
More in Comparisons
- Novita AI API alternative for video and image jobs: Sume
Novita offers model APIs, agent sandboxes and GPU deployment. Sume offers managed media jobs only. Where they overlap and where Novita does more.
- OpenRouter models fallback array and 3-entry limit vs Sume
OpenRouter's models array tries the next model on downtime, rate limits or moderation; fallbacks allows 3. Sume's allow_fallbacks has no effect.
- OpenRouter provider.sort and max_price vs Sume's inert sort
OpenRouter's provider.sort picks price, throughput or latency and turns off load balancing. On Sume's image route, sort is accepted and changes nothing.
- Pexels free stock footage vs generated B-roll: what to use when
Compare Pexels licensed footage with generated B-roll: licence limits, control, cost and where each wins, with Pexels terms to re-check.
Written by Sume