Grok Image on Sume: one image per call, so fan out four in Python
x-ai/grok-image lists n as 1 to 1 in the Sume catalog. How to get four variants with four parallel calls, what it costs, and how to stay under queue limits.

On Sume, x-ai/grok-image takes exactly one image per request: its n descriptor in the image catalog is a range of 1 to 1. To get four variants, send four requests in parallel and collect the four results; one request with n: 4 returns 400 invalid_request with the message that n must be between 1 and 1 for that model.
That is a Sume catalog fact, not a general statement about the model. The Image API docs say n goes up to 10 per call but that per-model ceilings are lower, and tell you to read the n range from the catalog. Every other model in the catalog I read on 2026-10-10 allows up to 4, except Higgsfield Soul, which allows only 1 or 4.
What the catalog says about n
The table lists the image catalog as read on 2026-10-10. Prices are the catalog's base price for one image; quality, resolution and other settings can change what a given request costs, so read the endpoint pricing for your exact settings.
| Model | n allowed | Base price per image (USD) |
|---|---|---|
| x-ai/grok-image | 1 | 0.025 |
| qwen/qwen-image | 1 to 4 | 0.025 |
| google/imagen-4-fast | 1 to 4 | 0.025 |
| black-forest-labs/flux.2-pro | 1 to 4 | 0.0375 |
| higgsfield/soul | 1 or 4 only | 0.005 |
Four calls instead of n=4
Billing works per image: you pay cost_usd times n, and a failed or cancelled generation is not billed. Four separate one-image calls therefore cost the same as one n: 4 call would, 4 x 0.025 = 0.10 USD at the base price. You lose only the convenience of one response.
The cost of fan-out is concurrency. Sume accepts valid jobs as queued while queue capacity remains, but when the workspace cannot take another paid job it returns 429 queue_full, and a plain 429 rate_limited means your request rate is too high. The Errors and rate limits page says to back off, to use retry-after when it is present, and not to retry unsafe submits without an Idempotency-Key. A semaphore keeps the burst small.
import asyncio, os
import httpx
URL = "https://api.sume.com/v1/images"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
async def one(client, sem, prompt):
async with sem:
body = {"model": "x-ai/grok-image", "prompt": prompt, "n": 1}
r = await client.post(URL, headers=HEADERS, json=body)
r.raise_for_status()
return r.status_code, r.json()
async def main():
sem = asyncio.Semaphore(4)
prompts = [f"a red scooter on a studio floor, take {i}" for i in range(1, 5)]
async with httpx.AsyncClient(timeout=60) as client:
out = await asyncio.gather(*(one(client, sem, p) for p in prompts))
for code, data in out:
print(code, data["data"][0]["url"] if code == 200 else data["data"]["status_url"])
asyncio.run(main())What Grok Image still offers
One image per call does not make the model a poor fit. In the same catalog Grok Image accepts up to 10 input_references, so it can edit a photo, and its aspect_ratio list includes tall phone shapes: 2:1, 20:9, 19.5:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, 9:19.5, 9:20 and 1:2. If you need a tall phone-screen frame such as 9:20 or 19.5:9, that list is a reason to choose it.
At 0.025 USD base price it sits in the lowest price group of the catalog, level with Qwen Image and Imagen 4 Fast, so four takes cost ten cents. The one thing to plan for is code: a loop of one-image requests, written once, is the whole difference.
Handle the 202 inside the fan-out
Each of the four calls can independently return 200 with the image or 202 with a job envelope, as the Image API docs describe. The sample prints the status URL for a 202, but a real pipeline should poll it with exponential backoff and stop on completed, failed or canceled, as Jobs and results advises.
Do not resubmit a call because one of the four is slow. The other three are done, and a resubmit would create and bill a fifth job.
- Put the variation in the prompt text, because
seedis not served and returns400 unsupported_parameter. - Keep the semaphore at or below the number of variants you want at once.
- If you need
n: 4in one call at about the same price, switch to Qwen Image or Imagen 4 Fast and compare the results yourself.
Sources
Related posts
More in Developers
- Grok Imagine ignores aspect_ratio on image-to-video; Sume rejects it
xAI says image-to-video output matches the input image and ignores aspect_ratio. Sume's Grok row goes further and rejects the field. Crop the still first.
- Grok Imagine Lite's 10 requests per second vs Sume plan concurrency
xAI lists a 10 requests per second limit and Batch API for Grok Imagine 1.5 Lite. On Sume, a clip batch is bounded by plan concurrency and queue capacity.
- How many Format runs per minute can my Sume plan start?
Sume limits writes per minute by plan, from 120 on Free to 1200 on Scale, with reads at 40 times that. Read the rate-limit headers and back off on 429.
- HyperFrames check via the Sume API: caption collisions pre-render
Send check with caption_zone to POST /v1/hyperframes-previews and get findings, contrast and overlap reports. A failing check is still a completed job.
Written by Sume