Nano Banana Pro interleaved text and images vs Sume image output

Google documents Nano Banana Pro returning text blocks with illustrations in one answer. Sume's Image API documents image results only; here is the workaround.

5 min readSume
All posts

Can you get a story or a how-to guide with text and pictures woven together from Nano Banana Pro through Sume? Not in one call. Google documents interleaved output for Nano Banana Pro, but Sume's Image API is documented as a prompt-in, images-out route, so you build the text-and-picture sequence yourself, one image job per step.

This post separates the two claims. Google's side comes from the Gemini API image generation page, read on 2026-10-03. Sume's side comes from the Image API docs and Jobs and results. Where the Sume docs are silent we say so rather than guess.

What does Google say Nano Banana Pro can interleave?

Google's page describes Nano Banana Pro (gemini-3-pro-image) as the premium tier of the family, and says Gemini 3 Pro can generate interleaved content, such as stories or instructional guides that contain both text blocks and illustrations. In that mode a single model answer carries several parts: a paragraph, an image, the next paragraph, the next image.

The same page lists the other behaviours that shape that kind of output. Thinking is enabled by default for the Gemini 3 image models and cannot be disabled, and the model can produce up to two interim images while it tests a composition. Nano Banana 2 and Pro offer 512px, 1K, 2K and 4K output, and every generated image carries a SynthID watermark.

What does Sume's Image API return?

The Image API takes POST /v1/images with a model and a prompt, and optional reference images. A finished request answers with data[], and each entry holds a Sume-hosted url and a media_type. A slower request answers 202 with a job envelope, and the images are then read from GET /v1/jobs/{id}/result as result.artifacts[] entries of type image.

Neither shape documents a text field. The docs describe the route as generating images from text prompts and reference images, and the usage object reports cost with token counts fixed at 0. So the part of Google's feature that returns prose is not a documented part of the Sume contract. We cannot show you a Sume field that carries the model's accompanying text, and this post does not claim one exists.

How do you build the sequence on Sume instead?

Split the work the way a layout tool would. You write or generate the copy yourself, then ask for one image per step, and you assemble the pairs in your own page, document or deck.

A practical pattern keeps the visual style consistent: generate the first image, then pass it as a public HTTPS reference to the later steps with a short instruction such as keep the same character and palette. Sume's reference-image rules are in the Image API docs: references must be public HTTPS URLs, and a model whose input_references descriptor is {min: 0, max: 0} rejects them.

  • Write the step list first: caption text, then one image prompt per caption.
  • Send each image prompt as its own request with its own Idempotency-Key, so a retry after a client timeout cannot bill twice.
  • Pass the first approved image as a reference to later steps to hold style.
  • Store the Sume media URLs from the result, not any provider URL.

How do you run many steps without a loop of your own?

On hosted MCP, the MCP tools and gates page describes script_run, a short JavaScript program that runs on the Sume side and can call generate_image once per scene. It fits the case of three or more calls of the same shape. Each paid call still needs its own idempotency_key, the run is bounded by timeout_seconds (5 to 55), max_calls and max_paid_calls, and the answer carries the child jobs[] for a batch jobs_wait.

Over plain HTTP the same idea is a loop that submits N jobs with mode: "async" and polls their status endpoints. Either way each picture is a separate billed job.

Which approach fits which output?

Before building, read the model's capability record with GET /v1/images/models; a request that sets a parameter the model does not list is rejected with 400 unsupported_parameter. Pricing for each billable line is on the endpoint record, and a failed or cancelled generation is not billed.

Interleaved answer versus one image per step (read 2026-10-03)
NeedGoogle's documented Nano Banana Pro behaviourWhat to do on Sume
Text and pictures in one answerInterleaved content for stories and guidesWrite the text yourself, one image job per step
Same look across stepsModel keeps context in one answerPass the first image as a public HTTPS reference
Predictable costNot covered on the pageRead the per-image pricing line from the model's endpoint record
Progress for a slow 4K stepThinking can add interim imagesUse mode: async and read job status

What should you not expect?

Do not expect the model's reasoning images to appear as separate outputs, and do not expect captions to come back with the pictures. If you need an author for the captions, write them yourself or draft them in a separate text step first. The narrow claim here is the documented one: Google describes interleaved output, and Sume's documented Image API result is images.

Sources

Related posts

More in Models

All Models posts

Written by Sume