Sume image models at a $0.02 list price: Grok, Qwen, Imagen Fast
Three Sume image models list at $0.02 per image: Grok Imagine, Qwen Image and Imagen 4 Fast. How they differ on edits, ratios and image count per call.

Three image models in Sume's catalog share the lowest standard list price of $0.02 per image: Grok Imagine (grok-image), Qwen Image (qwen-image) and Imagen 4 Fast (imagen-4-fast). At the catalog's list times 1.25 rule that is $0.025 per image. They are not interchangeable: only two take edits, only one is limited to a single image per call, and their aspect ratio lists differ.
The numbers below come from the image router catalog and price table in Sume's repo, read 2026-10-03. Billing is the catalog formula, list times 1.25, so check GET /v1/images/models for the live amount before you budget.
The three side by side
Same list price, different capabilities. If you only need a cheap text-to-image draft, any of the three works. If you need to edit or reference an existing picture, Imagen 4 Fast is out.
| Id | Edit with image_urls | Max images per call | Aspect ratios listed |
|---|---|---|---|
| grok-image | Yes | 1 | 13 (2:1, 20:9, 19.5:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, 9:19.5, 9:20, 1:2) |
| qwen-image | Yes | 4 | 13 (includes 21:9, 9:21, 4:5, 5:4) |
| imagen-4-fast | No (text only) | 4 | 5 (1:1, 16:9, 9:16, 4:3, 3:4) |
Which one for which job
Grok Imagine is the only one of the three that lists tall phone shapes such as 9:20 and 9:19.5, and it returns one image per call, so a four-way variant grid costs four calls. Qwen Image lists 21:9 and 9:21 plus 4:5, which covers Instagram portrait and ultrawide banners, and it can return up to four images in one call. Imagen 4 Fast is the plainest: five ratios, text prompts only, up to four images.
A rule of thumb that follows from the table:
- Need a 4:5 feed image? Qwen Image lists it; Grok Imagine and Imagen 4 Fast do not.
- Need a very tall wallpaper? Grok Imagine lists 9:20 and 9:19.5.
- Need to restyle a photo you already have? Grok Imagine or Qwen Image, not Imagen 4 Fast.
- Need four options in one call? Qwen Image or Imagen 4 Fast.
What xAI says about its own tiers
xAI's models page (read 2026-10-03) lists three Grok image ids: grok-imagine-image at $0.02 per image, grok-imagine-image-2.0 at $0.04 and grok-imagine-image-quality at $0.05. Sume's catalog row is the plain grok-image id priced at $0.02 list, which matches the first of those three. The page does not state whether Sume resells the others, and Sume's catalog lists one Grok image row, so treat Grok Imagine Image 2.0 as not on Sume until GET /v1/images/models shows it.
Sume's docs also say a request that sets a parameter the model does not list is rejected with 400 unsupported_parameter, so a mask_url or a reference on Imagen 4 Fast fails fast instead of being ignored.
Read the live list first
List prices change and rows are added. Before you pin a model in a pipeline, call GET /v1/images/models and read each row's supported_parameters and pricing; the Image API page explains the descriptors. For a wider cost ladder, see image models under 4 cents.
Sources
Related posts
More in Models
- Veda sparse attention for MiniMax H3: 6.8x attention, 3.1x clip
Veda's sparse attention keeps 10 percent of attention work for MiniMax H3. Why 6.8x on attention becomes 3.1x per clip, and what it needs to run.
- Gemini 3.8 Flash is stable: keep model ids in config
When a vendor ships a new Flash model, a hard-coded id ages. Read Sume ids from GET /v1/catalog and treat 404 model_not_found as a signal, not a retry.
- How long can a Veo 3.1 video get? 148 seconds at 720p
Veo 3.1 extends clips 7 seconds at a time, up to 148 seconds at 720p. Sume has no extend task, so here are the longer single-clip models and how to join clips.
- Veo 3.1 Lite 1080p is 8 seconds only: options for a 5-second clip
Veo 3.1 Lite renders 1080p only at 8 seconds. To get a 5-second 1080p clip, trim an 8-second render with Sume or pick a model whose range includes 5.
Written by Sume