Nano Banana 2.1 replies with text and images; Sume returns URLs only
Google says Nano Banana 2.1 outputs text as well as images. Sume's image response lists image URLs and usage, so write captions and alt text in a separate step.

Nano Banana 2.1 can answer with text as well as pictures when you call it through Google, but the image response you get from Sume contains image URLs, a media type and a usage block, and no model text. If your app depends on a caption, an explanation or alt text coming back with the image, you need a second step that you own.
The Google side is from the Nano Banana 2.1 model card and the Gemini API image generation guide, read 2026-10-10. The Sume side is the documented response in the Image API.
What Google says the model returns
The model card describes Nano Banana 2.1 as taking text and images as input and producing images and text as output, with a 64K text output token budget alongside 4K image output tokens and a 1M-token input context. The API guide lists multi-turn conversational editing for every model in the family, and interleaved text and images for Gemini 3 Pro Image specifically. So text in the response is a documented behavior of the Google API, and the guide treats it as normal for editing conversations.
What Sume's response contains
A successful POST /v1/images returns a created timestamp, the model you asked for, a data array and a usage object. Each data item has a url on media.sume.com and a media_type. The usage block carries cost in US dollars, and the docs state token counts are always 0 because Sume meters image models per image. There is no field for text the model wrote alongside the picture.
A slow call returns 202 with a job envelope instead; the result of that job follows the standard job result shape in Jobs and results, not the image body. Either way, the contract is images and cost.
| Output | Google API (Nano Banana 2.1) | Sume POST /v1/images |
|---|---|---|
| Image | Yes, up to 4K output | Yes, hosted URL per item |
| Text in the reply | Yes, per the model card | No field in the documented response |
| Multi-turn state | Conversational editing in the guide | None; each call is a fresh request |
| Cost reporting | Not stated on the card | usage.cost in USD per call |
Three ways to get the words you need
Keep the Sume job metadata so the prompt, model and result stay together for whoever edits it next.
- Put the copy in the prompt. If you want a label or headline inside the picture, say the exact string in the prompt; that needs no extra text channel.
- Write captions yourself. Generate the image on Sume, then write alt text or a caption in your own code or tool, using the prompt you sent as the starting point.
- Pass the previous result back as a reference. For a follow-up edit, send the earlier image URL in
input_referenceswith a fresh instruction; Sume does not keep a conversation, so each edit carries its own full instruction.
Why it matters for edit loops
A common Google flow is to ask for an edit, read the model's note about what it changed, and then ask again. On Sume the note never arrives, so build the loop around inspection: look at the image, restate what to keep, and name only the change. Add the instruction to keep everything else identical, and use aspect_ratio: "auto" on edit calls so the output matches the reference shape; omitting the field is not the same as auto, per the Image API docs.
If you need descriptions at scale, such as alt text for a 30-image gallery, plan it as its own pass after the images exist, with the image URL and the original prompt as input, and store the result next to the job. That also keeps image spend and text spend separate on the bill, which makes both easier to audit.
None of this makes Sume a worse place to generate pictures. It means the contract is narrower than the Google chat surface, and it is better to know that before you design a feature on the assumption of a text reply.
Sources
Related posts
More in Models
- Qwen Image 2.1 Turbo at $0.016 an image vs Qwen Image on Sume
Qwen Image 2.1 Turbo's API launched Oct 9 at $0.016 an image in Singapore. Sume does not list it; its Qwen Image row works out to $0.025 before rounding.
- Seedance 2.0 11-second clip price on Sume, with other options
A 11-second Seedance 2.0 clip bills $1.9335 at 480p and $9.3555 at 1080p on Sume. Other models at 1080p for comparison.
- Seedance 2.0 12-second clip price on Sume, with other options
A 12-second Seedance 2.0 clip bills $2.1092 at 480p and $10.206 at 1080p on Sume. Other models at 1080p for comparison.
- Seedance 2.0 13-second clip price on Sume, with other options
A 13-second Seedance 2.0 clip bills $2.285 at 480p and $11.0565 at 1080p on Sume. Other models at 1080p for comparison.
Written by Sume