Nano Banana text in images: write the copy first, then render
Google's tip for text in Gemini images: settle the wording first, then ask for the image. How to do that with one Sume request and a short checklist.

To get readable text in a Nano Banana image, decide the exact wording before you ask for the picture, then put that wording in the prompt in quotes. Google's Gemini image docs say that when generating text for an image, Gemini works best if you first generate the text and then ask for an image with the text. On Sume there is no separate text step: you write the copy yourself, or ask a chat model for it, and send the final string inside prompt.
The tip is from Google's Image generation with Gemini page; the request shape is from the Sume Image API page. Both were read on 2026-10-02.
What does Google actually recommend?
The Gemini page's limitations section says: when generating text for an image, Gemini works best if you first generate the text and then ask for an image with the text. It also says the model will not always follow the exact number of images you ask for, which matters if you ask for a set of posters with different slogans in one call.
What does that look like as a Sume call?
Fix the wording first, then send it verbatim. The Sume Image API takes a required plain-text prompt; the request below uses the catalog id google/nano-banana-2. Sume returns 200 with data[].url when the image finishes inside its 30-second blocking budget, and 202 with a job envelope when it does not.
import os, requests
body = {
"model": "google/nano-banana-2",
"prompt": ("Poster for a bakery. The headline reads exactly: "
"FRESH BREAD DAILY. Warm morning light, flour dust, "
"bold sans-serif headline at the top."),
}
r = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json=body,
timeout=60,
)
print(r.status_code)
print(r.json())How should I handle several slogans?
Send one request per slogan instead of asking for several in one prompt. Google says the model may not follow an exact output count, and Sume bills each completed generation per image, so a failed or wrong slogan costs one retry, not a whole batch. Sume's docs say failed or cancelled generations are not billed.
| Step | Do this | Source of the rule |
|---|---|---|
| 1 | Write the final copy as a string | Google: generate the text first |
| 2 | Quote it in the prompt and say where it goes | Plain prompt field on Sume |
| 3 | One slogan per request | Google: exact image counts not guaranteed |
| 4 | Check the spelling by eye before publishing | Failed generations are not billed on Sume; wrong text still is |
What if the text still comes out wrong?
A completed image is billed even if a letter is wrong, so re-check before fanning out. Shorten the copy, name the font style in words, and re-run. If a wrong word is the only problem, edit the image with the corrected wording as a reference prompt rather than starting over.
Do not assume a higher-priced model fixes spelling. The Sume docs publish per-image prices for each model on the catalog endpoint; compare them there before moving up.
Sources
Related posts
More in Models
- OpenAI media model shutdown calendar 2026: DALL-E, Sora, GPT Image
Three OpenAI media retirements in 2026: DALL-E on May 12, Sora and the Videos API on September 24, GPT Image 1.5 and mini on December 1. Dates and replacements.
- Qwen-Image 2.0 Pro is Alibaba's pick: which Qwen ids does Sume list?
Alibaba recommends qwen-image-2.0-pro. Sume's image catalog lists qwen/qwen-image and qwen/qwen-image-max in the repo; check the live catalog for more.
- Qwen-Image negative_prompt 500 characters vs a Sume image request
Alibaba's Qwen-Image API takes a negative_prompt up to 500 characters. Sume's image request table has no such field, so rewrite exclusions into the prompt.
- Seedance 2.5 concert prompt uses 18 images: fitting Sume's 9
Seed's Seedance 2.5 concert example uses 18 reference images. Sume's Video Router takes at most 9 for seedance-2.5, so merge groups into sheets.
Written by Sume