Imagen 4 Fast vs Ultra on Sume: $0.025 vs $0.075, text-to-image only
Imagen 4 Fast bills $0.025 and Ultra $0.075 on Sume. Both are text-to-image only with five ratios and no 4:5. Ultra adds 1K and 2K tiers. What that means.

Imagen 4 on Sume comes in two rows: Fast at $0.025 per image and Ultra at $0.075, both billed as list rate times 1.25. Neither accepts reference images, so they cannot edit a photo, and neither lists 4:5. If you need either of those, use another model.
The two rows
| Field | Imagen 4 Fast | Imagen 4 Ultra |
|---|---|---|
| API id | google/imagen-4-fast | google/imagen-4-ultra |
| List per image | $0.02 | $0.06 |
| Billed (list x 1.25) | $0.025 | $0.075 |
| Reference images | No (text-to-image only) | No (text-to-image only) |
| Aspect ratios | 1:1, 16:9, 9:16, 4:3, 3:4 | 1:1, 16:9, 9:16, 4:3, 3:4 |
| Resolution tiers listed | None | 1K, 2K |
| Max n per call | 4 | 4 |
What "no 4:5" means for social
Instagram portrait is 4:5 (1080x1350). Imagen's list stops at 3:4, so a feed image for that slot would be a different shape from what the platform shows. The Image API notes call this out: Imagen and Grok do not include 4:5, while Nano Banana, Seedream, Flux, Qwen and Ideogram do. Sume does not map 4:5 to 4:3 or 3:4 for you, because they are different shapes.
For 9:16 covers and 16:9 thumbnails, both Imagen rows are fine on shape.
Fast for drafts, Ultra for finals?
That is the usual way to use a two-tier family, and the numbers support the idea of drafting: 40 Fast drafts cost 40 x $0.025 = $1.00, which is the price of 13.3 Ultra images. Finalise only the prompts that earned it, at $0.075 each. Sume does not claim a specific quality gap, so look at your own outputs before committing a pipeline to it.
Ultra's 1K and 2K tiers are in its catalog row. The repo's rate table carries one Ultra rate, so a tier change should not surprise you on price, but read the endpoint's pricing line before a large run.
Other rows if you hit the limits
When you need a reference image as input, Imagen cannot help: Imagen 4 Fast and Ultra, Recraft V4, Qwen Image Max and Soul are the text-only rows in the catalog. Read supported_parameters for the row before you plan an edit workflow.
Price is the reason to stay on Imagen Fast: at $0.025 it sits with Grok Imagine and Qwen Image as the cheapest standard rows, with only Soul lower. If a text-to-image prompt works there, you do not need a dearer row.
Avoid the 400
Sending input_references to either row returns 400 unsupported_parameter, as does mask_url or background. Check the descriptor (input_references with min 0 and max 0 means text-only) in GET /v1/images/models first, as the Image API page describes.
Sources
Related posts
More in Models
- Imagen 4 shut down Aug 17: Nano Banana 2.1 is the named replacement
Google lists gemini-nano-banana-2.1 as the Imagen 4 replacement. Imagen rows on Sume take no reference photos, while Nano Banana 2.1 takes up to 10.
- Lip-sync and motion control cost per clip: H3 Max, Fabric, Kling
Sume prices H3 Max lip-sync at $0.0625 to $0.20 per audio second, Fabric at $0.1875 and Kling 3.0 Motion Control at $0.1575. A 10-second clip is $0.63 to $2.00.
- Music Router prompt limit: 5,000 characters, and what to put in them
Sume Music Router prompts run 1 to 5,000 characters, with no duration field or negative prompt; length goes in the text. A brief budget and a checker.
- MAI-Transcribe-2-Streaming 0.13 s to final: when the clock starts
The 0.13 second figure is measured from end of speech found by a VAD, and partials arrive in about 100 ms. Why neither number is the wait for a file transcript.
Written by Sume