200 drafts at 2.5 cents: Imagen 4 Fast, Qwen Image or Grok Image
Each costs $5.00 for 200 images on Sume. Imagen 4 Fast and Qwen Image take n up to 4, so 50 calls; Grok Image takes n of 1, so 200. Ratios and edits differ.

Imagen 4 Fast, Qwen Image and Grok Image all bill $0.025 per image on Sume, so 200 draft images cost $5.00 on any of them. The call count is what differs: Imagen 4 Fast and Qwen Image accept n up to 4, so 200 images take 50 calls, while Grok Image accepts only n of 1, so the same batch takes 200. Only Qwen Image and Grok Image accept reference images; Imagen 4 Fast is text-to-image only.
The figures come from the Sume catalog on origin/main, read 2026-10-08. Charge per call is cost_usd × n, so splitting a batch changes the number of requests, not the price.
The three rows
Aspect lists are the main content difference. Imagen has five ratios. Qwen Image has thirteen, including 21:9 and 9:21. Grok Image has thirteen too, but its set is phone-shaped: 20:9, 19.5:9, 9:19.5 and 9:20 appear, and 21:9 does not.
| Field | Imagen 4 Fast | Qwen Image | Grok Image |
|---|---|---|---|
| Billed per image | $0.025 | $0.025 | $0.025 |
| 200 images | $5.00 | $5.00 | $5.00 |
| n per call | 1 to 4 | 1 to 4 | 1 only |
| Calls for 200 | 50 | 50 | 200 |
| Reference images | None | Up to 10 | Up to 10 |
| Aspect ratios | 5 | 13, with 21:9 | 13, with 20:9 and 19.5:9 |
| Output formats | png, jpeg, webp | png, jpeg, webp | png, jpeg, webp |
What 200 calls means
A call that finishes inside 30 seconds returns 200 with the images. One that does not returns 202 and a job envelope, and you poll status_url as the jobs page describes. A draft batch of 200 single-image calls multiplies that polling, so send Grok Image batches in modest waves and respect next_poll_after_seconds when it is present.
The 50-call routes have the easier loop. If the content of the draft does not need a specific ratio, Imagen 4 Fast in fours is the least code to run.
Choosing among them
Need a source photo in the request: Qwen Image or Grok Image. Need an ultrawide: Qwen Image. Need a phone-tall 20:9 or 9:20 draft: Grok Image. Need only squares and 16:9 from text: Imagen 4 Fast.
None of these rows lists quality or resolution, and none accepts seed, so repeats are fresh samples.
Write the batch runner so that the unit of work is one call, and set the number of images per call from the row's n maximum, which you can read from the catalog. That way a row with a different ceiling needs no code change, and a Grok Image run is the same loop with more iterations.
Sources
Related posts
More in Comparisons
- First paid AI avatar plan, October 2026: $18 to $59 vs Sume seconds
Cheapest paid plans at Synthesia, HeyGen, Colossyan and Tavus from their pricing pages, with the seconds of Sume Standard avatar video the same money buys.
- AI avatar news, October 8, 2026: what is worth acting on
Synthesia Sessions and plans, Tavus Griffin-Lite, HeyGen LiveAvatar, Teams and Zoom avatars: dated, with what a buyer can actually do this week.
- Amazon no-text lower-right rule vs Walmart contrast: where captions go
Amazon Sponsored Brands bans text in the lower right; Walmart wants 4.5:1 contrast. How to move Sume burned-in captions up with anchor ratios and check a still.
- Claude MCP connector or Sume Agent Completions for a long video job?
A Messages API request with the MCP connector holds open while Claude works. Sume Agent Completions return 202 and a run to poll. Which one fits a long job.
Written by Sume