Packaging label text test: 12 prompts on 3 Sume models, what it costs

Test Nano Banana 2.1, Imagen 4 Ultra and GPT Image 2.5 on 12 label-text prompts first: about $1.20, $0.90 and $1.40 on Sume. A short script runs it.

5 min readSume
All posts

Before you pick a model for packaging text, run the same 12 prompts through each candidate and read the output yourself. On Sume that costs about $1.20 on Nano Banana 2.1 at 1K, $0.90 on Imagen 4 Ultra and about $1.40 on ChatGPT Image 2.5 at xhigh, so the whole test is under $4. Google's own pricing page says Nano Banana 2.1 is built for "accurate text rendering", but that is a vendor claim; this test is how you check it for your label.

What each test costs

Nano Banana 2.1 bills $0.10 at 1K on Sume. Imagen 4 Ultra bills $0.075. For GPT Image 2.5 the Sume docs give the provider output estimate at 1024x1024 as $0.09366 for xhigh and $0.21072 for max, before input tokens and Sume pricing. Times 1.25 that is about $0.117 for xhigh, output only, so the 12-prompt figure is an estimate that excludes input tokens.

12-prompt label-text test, Sume pricing as of 2026-10-08
Model idPer imageArithmetic12 prompts
google/nano-banana-2.1 (1K)$0.1012 x 0.10$1.20
google/imagen-4-ultra$0.07512 x 0.075$0.90
openai/gpt-image-2.5 (xhigh, 1024x1024)about $0.11712 x 0.117075about $1.40 plus input tokens

Write prompts that can fail

Use real label copy: a brand name with an unusual spelling, a net-weight line like "NET WT 12 OZ (340 g)", a small ingredient list, and one line in a second language if your market needs it. Score each output on spelling, line breaks and whether small print is readable. Do not score on how nice the bottle looks.

The script

This sends each prompt to each model and prints the status and the first URL. quality is only valid on the GPT row, because a model that does not list a parameter returns 400 unsupported_parameter.

import os, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODELS = {
    "google/nano-banana-2.1": {"resolution": "1K"},
    "google/imagen-4-ultra": {},
    "openai/gpt-image-2.5": {"quality": "xhigh", "image_size": "1024x1024"},
}
PROMPTS = [
    'Cardboard coffee bag, label reads "KAFFEEHAUS NORD" and "NET WT 12 OZ (340 g)"',
    'Jar of chili crisp, front label "HOT & SMOKY" with a 4-line ingredient list',
]
for model, extra in MODELS.items():
    for p in PROMPTS:
        body = {"model": model, "prompt": p, **extra}
        r = requests.post("https://api.sume.com/v1/images", headers=H, json=body, timeout=60)
        j = r.json()
        if r.status_code == 200:
            print(model, 200, j["data"][0]["url"])
        else:
            print(model, r.status_code, j)

Read the result

A 200 response carries the image. A 202 means the job outlived the 30-second budget and you poll the status URL; see Jobs and results. Keep the model with the fewest text errors that fits your price, then re-run only that model on your real catalog.

Reading the scores

Give each output a score from 0 to 3 for each prompt: 0 for unreadable, 1 for one or more wrong characters, 2 for correct text but poor layout, 3 for correct and usable. Add the scores per model. With 12 prompts the maximum is 36, and a gap of three or four points is small enough to be noise, so run a second batch of 12 before you act on it.

Repeat the test at the size you will ship. A label that reads well at 1024x1024 can fall apart when the final is 4K and you zoom in on a small ingredient line. If the winner is Nano Banana 2.1, price the extra tiers: 2K is $0.15 and 4K is $0.20 per image, so a 12-prompt re-run at 2K costs 12 x $0.15 = $1.80.

Keep the raw prompts and the returned URLs in a file. When the catalog changes, you can re-run the same 12 prompts on a new row and compare against your earlier scores.

Sources

Related posts

More in Models

All Models posts

Written by Sume