Nano Banana Pro interleaved text and images vs Sume image output
Google documents Nano Banana Pro returning text blocks with illustrations in one answer. Sume's Image API documents image results only; here is the workaround.

Can you get a story or a how-to guide with text and pictures woven together from Nano Banana Pro through Sume? Not in one call. Google documents interleaved output for Nano Banana Pro, but Sume's Image API is documented as a prompt-in, images-out route, so you build the text-and-picture sequence yourself, one image job per step.
This post separates the two claims. Google's side comes from the Gemini API image generation page, read on 2026-10-03. Sume's side comes from the Image API docs and Jobs and results. Where the Sume docs are silent we say so rather than guess.
What does Google say Nano Banana Pro can interleave?
Google's page describes Nano Banana Pro (gemini-3-pro-image) as the premium tier of the family, and says Gemini 3 Pro can generate interleaved content, such as stories or instructional guides that contain both text blocks and illustrations. In that mode a single model answer carries several parts: a paragraph, an image, the next paragraph, the next image.
The same page lists the other behaviours that shape that kind of output. Thinking is enabled by default for the Gemini 3 image models and cannot be disabled, and the model can produce up to two interim images while it tests a composition. Nano Banana 2 and Pro offer 512px, 1K, 2K and 4K output, and every generated image carries a SynthID watermark.
What does Sume's Image API return?
The Image API takes POST /v1/images with a model and a prompt, and optional reference images. A finished request answers with data[], and each entry holds a Sume-hosted url and a media_type. A slower request answers 202 with a job envelope, and the images are then read from GET /v1/jobs/{id}/result as result.artifacts[] entries of type image.
Neither shape documents a text field. The docs describe the route as generating images from text prompts and reference images, and the usage object reports cost with token counts fixed at 0. So the part of Google's feature that returns prose is not a documented part of the Sume contract. We cannot show you a Sume field that carries the model's accompanying text, and this post does not claim one exists.
How do you build the sequence on Sume instead?
Split the work the way a layout tool would. You write or generate the copy yourself, then ask for one image per step, and you assemble the pairs in your own page, document or deck.
A practical pattern keeps the visual style consistent: generate the first image, then pass it as a public HTTPS reference to the later steps with a short instruction such as keep the same character and palette. Sume's reference-image rules are in the Image API docs: references must be public HTTPS URLs, and a model whose input_references descriptor is {min: 0, max: 0} rejects them.
- Write the step list first: caption text, then one image prompt per caption.
- Send each image prompt as its own request with its own
Idempotency-Key, so a retry after a client timeout cannot bill twice. - Pass the first approved image as a reference to later steps to hold style.
- Store the Sume media URLs from the result, not any provider URL.
How do you run many steps without a loop of your own?
On hosted MCP, the MCP tools and gates page describes script_run, a short JavaScript program that runs on the Sume side and can call generate_image once per scene. It fits the case of three or more calls of the same shape. Each paid call still needs its own idempotency_key, the run is bounded by timeout_seconds (5 to 55), max_calls and max_paid_calls, and the answer carries the child jobs[] for a batch jobs_wait.
Over plain HTTP the same idea is a loop that submits N jobs with mode: "async" and polls their status endpoints. Either way each picture is a separate billed job.
Which approach fits which output?
Before building, read the model's capability record with GET /v1/images/models; a request that sets a parameter the model does not list is rejected with 400 unsupported_parameter. Pricing for each billable line is on the endpoint record, and a failed or cancelled generation is not billed.
| Need | Google's documented Nano Banana Pro behaviour | What to do on Sume |
|---|---|---|
| Text and pictures in one answer | Interleaved content for stories and guides | Write the text yourself, one image job per step |
| Same look across steps | Model keeps context in one answer | Pass the first image as a public HTTPS reference |
| Predictable cost | Not covered on the page | Read the per-image pricing line from the model's endpoint record |
| Progress for a slow 4K step | Thinking can add interim images | Use mode: async and read job status |
What should you not expect?
Do not expect the model's reasoning images to appear as separate outputs, and do not expect captions to come back with the pictures. If you need an author for the captions, write them yourself or draft them in a separate text step first. The narrow claim here is the documented one: Google describes interleaved output, and Sume's documented Image API result is images.
Sources
Related posts
More in Models
- OmniVoice: 600+ languages, CC-BY-NC weights, hosted TTS instead
OmniVoice covers 600+ languages in a 0.6B model, but its weights are CC-BY-NC. What the card says, what it omits, and where a hosted TTS job fits.
- Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.
- Polish, Dutch, Swedish, Turkish text to speech API: Sume pl nl sv tr
Sume's Voices library tags voices pl, nl, sv and tr alongside 12 other languages. What to send for each, and how Eleven v4's list compares.
- QuantFunc INT4 MiniMax H3: 3.2 s per step on an RTX 4090
QuantFunc's 4-bit MiniMax H3 claims 3.2 s per step on an RTX 4090. That is not a clip time. What the card says, what it omits, and when to use a hosted job.
Written by Sume