Packaging label text test: 12 prompts on 3 Sume models, what it costs
Test Nano Banana 2.1, Imagen 4 Ultra and GPT Image 2.5 on 12 label-text prompts first: about $1.20, $0.90 and $1.40 on Sume. A short script runs it.

Before you pick a model for packaging text, run the same 12 prompts through each candidate and read the output yourself. On Sume that costs about $1.20 on Nano Banana 2.1 at 1K, $0.90 on Imagen 4 Ultra and about $1.40 on ChatGPT Image 2.5 at xhigh, so the whole test is under $4. Google's own pricing page says Nano Banana 2.1 is built for "accurate text rendering", but that is a vendor claim; this test is how you check it for your label.
What each test costs
Nano Banana 2.1 bills $0.10 at 1K on Sume. Imagen 4 Ultra bills $0.075. For GPT Image 2.5 the Sume docs give the provider output estimate at 1024x1024 as $0.09366 for xhigh and $0.21072 for max, before input tokens and Sume pricing. Times 1.25 that is about $0.117 for xhigh, output only, so the 12-prompt figure is an estimate that excludes input tokens.
| Model id | Per image | Arithmetic | 12 prompts |
|---|---|---|---|
| google/nano-banana-2.1 (1K) | $0.10 | 12 x 0.10 | $1.20 |
| google/imagen-4-ultra | $0.075 | 12 x 0.075 | $0.90 |
| openai/gpt-image-2.5 (xhigh, 1024x1024) | about $0.117 | 12 x 0.117075 | about $1.40 plus input tokens |
Write prompts that can fail
Use real label copy: a brand name with an unusual spelling, a net-weight line like "NET WT 12 OZ (340 g)", a small ingredient list, and one line in a second language if your market needs it. Score each output on spelling, line breaks and whether small print is readable. Do not score on how nice the bottle looks.
The script
This sends each prompt to each model and prints the status and the first URL. quality is only valid on the GPT row, because a model that does not list a parameter returns 400 unsupported_parameter.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODELS = {
"google/nano-banana-2.1": {"resolution": "1K"},
"google/imagen-4-ultra": {},
"openai/gpt-image-2.5": {"quality": "xhigh", "image_size": "1024x1024"},
}
PROMPTS = [
'Cardboard coffee bag, label reads "KAFFEEHAUS NORD" and "NET WT 12 OZ (340 g)"',
'Jar of chili crisp, front label "HOT & SMOKY" with a 4-line ingredient list',
]
for model, extra in MODELS.items():
for p in PROMPTS:
body = {"model": model, "prompt": p, **extra}
r = requests.post("https://api.sume.com/v1/images", headers=H, json=body, timeout=60)
j = r.json()
if r.status_code == 200:
print(model, 200, j["data"][0]["url"])
else:
print(model, r.status_code, j)Read the result
A 200 response carries the image. A 202 means the job outlived the 30-second budget and you poll the status URL; see Jobs and results. Keep the model with the fewest text errors that fits your price, then re-run only that model on your real catalog.
Reading the scores
Give each output a score from 0 to 3 for each prompt: 0 for unreadable, 1 for one or more wrong characters, 2 for correct text but poor layout, 3 for correct and usable. Add the scores per model. With 12 prompts the maximum is 36, and a gap of three or four points is small enough to be noise, so run a second batch of 12 before you act on it.
Repeat the test at the size you will ship. A label that reads well at 1024x1024 can fall apart when the final is 4K and you zoom in on a small ingredient line. If the winner is Nano Banana 2.1, price the extra tiers: 2K is $0.15 and 4K is $0.20 per image, so a 12-prompt re-run at 2K costs 12 x $0.15 = $1.80.
Keep the raw prompts and the returned URLs in a file. When the catalog changes, you can re-run the same 12 prompts on a new row and compare against your earlier scores.
Sources
Related posts
More in Models
- Pin lyria-3.5 or leave sume/music-auto? Read routed_model per job
sume/music-auto resolves to Lyria 3.5 today and can move without notice. When to pin lyria-3.5 on Sume and how job.request.routed_model records the engine.
- Premium image models on Sume ranked: 7 to 27 cents a picture
ChatGPT Image 2.5 high is 7 cents on Sume, Nano Banana 2.1 and Qwen Image Max 10, Nano Banana Pro 19, ChatGPT Image 2 27. Ranked, with caveats.
- 500-image job on a preview model: Hy Image 3.5 availability, plan B
OpenRouter shows 83.39% availability over 3 days for Hy Image 3.5 Preview, one provider. At 500 images that is about 83 failures to retry. Plan B on Sume.
- Product shots from one photo: 14 Sume edit models, 3 to 27 cents
A product-shot edit from one reference photo works on 14 Sume image models. 100 shots cost $3 to $27. Which also take 4:5, and when to use 16 references.
Written by Sume