Grok Image on Sume: one image per call, so fan out four in Python

x-ai/grok-image lists n as 1 to 1 in the Sume catalog. How to get four variants with four parallel calls, what it costs, and how to stay under queue limits.

4 min readSume
All posts

On Sume, x-ai/grok-image takes exactly one image per request: its n descriptor in the image catalog is a range of 1 to 1. To get four variants, send four requests in parallel and collect the four results; one request with n: 4 returns 400 invalid_request with the message that n must be between 1 and 1 for that model.

That is a Sume catalog fact, not a general statement about the model. The Image API docs say n goes up to 10 per call but that per-model ceilings are lower, and tell you to read the n range from the catalog. Every other model in the catalog I read on 2026-10-10 allows up to 4, except Higgsfield Soul, which allows only 1 or 4.

What the catalog says about n

The table lists the image catalog as read on 2026-10-10. Prices are the catalog's base price for one image; quality, resolution and other settings can change what a given request costs, so read the endpoint pricing for your exact settings.

n ceilings and base price per image in the Sume image catalog (read 2026-10-10)
Modeln allowedBase price per image (USD)
x-ai/grok-image10.025
qwen/qwen-image1 to 40.025
google/imagen-4-fast1 to 40.025
black-forest-labs/flux.2-pro1 to 40.0375
higgsfield/soul1 or 4 only0.005

Four calls instead of n=4

Billing works per image: you pay cost_usd times n, and a failed or cancelled generation is not billed. Four separate one-image calls therefore cost the same as one n: 4 call would, 4 x 0.025 = 0.10 USD at the base price. You lose only the convenience of one response.

The cost of fan-out is concurrency. Sume accepts valid jobs as queued while queue capacity remains, but when the workspace cannot take another paid job it returns 429 queue_full, and a plain 429 rate_limited means your request rate is too high. The Errors and rate limits page says to back off, to use retry-after when it is present, and not to retry unsafe submits without an Idempotency-Key. A semaphore keeps the burst small.

import asyncio, os
import httpx

URL = "https://api.sume.com/v1/images"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}


async def one(client, sem, prompt):
    async with sem:
        body = {"model": "x-ai/grok-image", "prompt": prompt, "n": 1}
        r = await client.post(URL, headers=HEADERS, json=body)
        r.raise_for_status()
        return r.status_code, r.json()


async def main():
    sem = asyncio.Semaphore(4)
    prompts = [f"a red scooter on a studio floor, take {i}" for i in range(1, 5)]
    async with httpx.AsyncClient(timeout=60) as client:
        out = await asyncio.gather(*(one(client, sem, p) for p in prompts))
    for code, data in out:
        print(code, data["data"][0]["url"] if code == 200 else data["data"]["status_url"])


asyncio.run(main())

What Grok Image still offers

One image per call does not make the model a poor fit. In the same catalog Grok Image accepts up to 10 input_references, so it can edit a photo, and its aspect_ratio list includes tall phone shapes: 2:1, 20:9, 19.5:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, 9:19.5, 9:20 and 1:2. If you need a tall phone-screen frame such as 9:20 or 19.5:9, that list is a reason to choose it.

At 0.025 USD base price it sits in the lowest price group of the catalog, level with Qwen Image and Imagen 4 Fast, so four takes cost ten cents. The one thing to plan for is code: a loop of one-image requests, written once, is the whole difference.

Handle the 202 inside the fan-out

Each of the four calls can independently return 200 with the image or 202 with a job envelope, as the Image API docs describe. The sample prints the status URL for a 202, but a real pipeline should poll it with exponential backoff and stop on completed, failed or canceled, as Jobs and results advises.

Do not resubmit a call because one of the four is slow. The other three are done, and a resubmit would create and bill a fifth job.

  • Put the variation in the prompt text, because seed is not served and returns 400 unsupported_parameter.
  • Keep the semaphore at or below the number of variants you want at once.
  • If you need n: 4 in one call at about the same price, switch to Qwen Image or Imagen 4 Fast and compare the results yourself.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume