Nano Banana 2 image search grounding: what Sume lets you send instead
Google lists Web Search and Image Search grounding for Nano Banana 2. Sume's image catalog has no grounding field, so fetch the references and pass up to 10.

Google's image generation docs say Nano Banana 2 (gemini-3.1-flash-image) integrates both Web Search and Image Search, and that Image Search can be used alone or together with Web Search (Google AI for Developers, read 2026-10-04). The same page lists up to 14 reference images in total. If you call Nano Banana 2 through Sume, you will look for the grounding switch and not find one.
What Sume exposes
Sume's image catalog publishes typed descriptors for each model: prompt, aspect ratio, resolution, quality, n, output format and input_references. A model accepts only the parameters it lists, and anything else returns 400 unsupported_parameter (Sume Image API). None of those descriptors is a search or grounding field, so there is nothing to switch on.
| Capability | Google API | Sume Image API |
|---|---|---|
| Web Search grounding | Listed | No field; 400 unsupported_parameter |
| Image Search grounding | Listed | No field; 400 unsupported_parameter |
| Reference images | Up to 14 in total | Up to 10 on this model per the Sume image-reference rule |
| Aspect ratio on edits | Fixed list of ratios | aspect_ratio auto matches the reference |
The workaround: be the search step
Grounding exists so the model can see what a real thing looks like. You can do that step yourself. Find the reference photos you are allowed to use, host them at public HTTPS URLs, and pass them in input_references. Say in the prompt what each image is. You decide which pictures count, and the job record shows exactly what was sent.
- Use only images you have the right to use as references.
- Keep to 10 references on Nano Banana 2 through Sume, per the docs' general rule of 10 (16 on ChatGPT Image 2.5).
- Log the reference URLs with the job, so a result can be traced to its inputs.
A request with fetched references
Post this body to POST /v1/images with your bearer key.
{
"model": "google/nano-banana-2",
"prompt": "A product shot of the object in Image 1, set in the kind of kitchen shown in Image 2. Keep the object's shape and label exactly.",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/object.png"}},
{"type": "image_url", "image_url": {"url": "https://example.com/kitchen.png"}}
],
"aspect_ratio": "auto"
}Sources
Related posts
More in Models
- Nano Banana reference image limits: Lite, 2 and Pro compared
Nano Banana 2 Lite takes up to 14 object images, Nano Banana 2 takes 10 object, 4 character and 3 style, Pro takes 6 object and 5 character.
- New speech models, October 2026: which ones a video maker can use
MAI-Voice-2.1, MAI-Transcribe-2-Streaming, Pocket TTS and Mercury Voice, sorted by what a short-video maker can run today, with the Sume step that matches each.
- One voice in 23 languages: Microsoft's claim vs Sume's language guard
Microsoft says MAI-Voice-2.1 uses one voice across 23 languages. Sume sets language per job and warns on a voice mismatch. What that means when you buy.
- Pixal3D multi-view: prepare the input images
Pixal3D added multi-view inference in September 2026 under an MIT license. Sume has no 3D, but reference-image edit can prepare input views.
Written by Sume