Grok Imagine returns one image per call: fan out 1,000 in Python
Grok Imagine on Sume takes n=1, so 1,000 images means 1,000 requests at $0.025 each, $25.00 in all. A short Python fan-out shows how to run them safely.

The Sume catalog lists Grok Imagine (x-ai/grok-image) with a maximum of one image per request, so a job of 1,000 images is 1,000 separate calls. At $0.025 per image, the bill is $25.00. Most other edit-capable models take up to four images per request, which would need only 250 calls for the same count.
That changes how you write the client. You cannot ask for a grid of variants in one call; you fan out, you cap the concurrency, and you total the cost from each response. A short Python version is below.
What the catalog says about n
The Image API docs describe n as an integer from 1 to 10 at the request level, with a per-model cap from the catalog. The source catalog sets Grok Imagine to one and most other rows to four. If you send a larger n than the model lists, the capability descriptor is the thing to check first: read n from GET /v1/images/models and do not hard-code the number.
Billing follows the images you receive. Sume bills list x 1.25 and you pay usage.cost for each response, so a call that returns one Grok image reports $0.025.
Requests needed for 1,000 images
The request count matters for rate and for time, not only for cost. The table compares Grok Imagine with a four-image model. The four-image row assumes the model accepts n: 4, which is the catalog default for the rows that do not state a lower cap.
| Model | Images per request | Requests | Billed per image | Total |
|---|---|---|---|---|
| Grok Imagine | 1 | 1,000 | $0.025 | $25.00 |
| Qwen Image | 4 | 250 | $0.025 | $25.00 |
| Flux 2 Pro | 4 | 250 | $0.0375 | $37.50 |
A fan-out that stays inside the rules
The script sends eight prompts four at a time, treats anything other than a 200 as not finished, and adds up usage.cost from the responses it got. It reads the key from SUME_API_KEY. A 202 means the job outlived the 30 second wait; a production version should then read the job result, as the jobs guide explains.
import json, os, urllib.error, urllib.request
from concurrent.futures import ThreadPoolExecutor
URL = "https://api.sume.com/v1/images"
KEY = os.environ["SUME_API_KEY"]
def one(prompt):
body = json.dumps({"model": "x-ai/grok-image", "prompt": prompt, "n": 1}).encode()
req = urllib.request.Request(URL, data=body, headers={
"Authorization": "Bearer " + KEY, "Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=60) as r:
data = json.load(r)
if r.status != 200:
return prompt, None, 0.0
return prompt, data["data"][0]["url"], data["usage"]["cost"]
except urllib.error.HTTPError:
return prompt, None, 0.0
prompts = [f"Product photo of a ceramic mug, variation {i}" for i in range(1, 9)]
with ThreadPoolExecutor(max_workers=4) as pool:
results = list(pool.map(one, prompts))
done = [r for r in results if r[1]]
print(len(done), "of", len(prompts), "images")
print("total", round(sum(r[2] for r in results), 4))
Why one image per call is not a defect
A one-image cap makes each call small and quick, which helps the 30 second synchronous wait: a single Grok image is unlikely to hit the 202 fallback, so your loop can treat a 200 as the normal path. It also makes failures cheap to reason about, because one failed request is one missing image, and a failed generation is not billed.
The cost is bookkeeping. Eight thousand requests need eight thousand log lines. If you want a grid of four variants from one prompt, a model with n up to 4 returns them in one response, and you pay the same total price per image; you only save on request overhead.
Limits to respect
- Start with a small worker count, such as four, and raise it only while errors stay at zero.
- Log each prompt with its URL and cost so a crash does not lose paid images.
- Stop when the running total passes your budget. Eight images should total $0.20, which is 8 x $0.025.
- Failed generations are not billed, so retry only the prompts without a URL.
Sources
Related posts
More in Developers
- Handle every Sume API error with one switch on next_action
Sume errors share one envelope. Branch on next_action, retryable and retry_after_seconds, and your client handles new codes without a code change. JS sample.
- HappyHorse 1.1's five aspect ratios vs the Sume video catalog
HappyHorse 1.1 on Cloudflare takes 16:9, 9:16, 1:1, 4:3 and 3:4. Which Sume video models cover all five, and which do not.
- HappyHorse 1.1 takes a seed on Cloudflare; Sume video models do not
Cloudflare's HappyHorse 1.1 accepts a seed up to 2,147,483,647 and an optional watermark. No Sume v1 video model accepts a seed; here is the workaround.
- Hindi speech to text API: Sume STT with language_code hi
Transcribe Hindi audio with Sume STT: send language_code hi, check the reported language, and review code-mixed speech. $0.01 per audio minute.
Written by Sume