Grok Imagine on Sume: n is 1, so four takes is four calls
Grok Imagine's image row on Sume caps n at 1 and costs $0.025. Four takes means four calls and $0.10. Other models accept n up to 4.

Answer
Grok Imagine on Sume (x-ai/grok-image) accepts n from 1 to 1, so one request returns one image and four takes cost four requests at $0.025 each, $0.10 in total. Nearly every other catalog row accepts n up to 4, which batches takes into one call at the same per-image price.
The n ceiling is a capability descriptor, not a guess, and it applies per request. GET /v1/images/models returns n: {type: range, min: 1, max: 1} for Grok Imagine. A request that sets a higher value is rejected rather than trimmed.
| Model | n range | Per image | 4 takes |
|---|---|---|---|
| Grok Imagine | 1 | $0.025 | $0.10, four calls |
| Qwen Image | 1 to 4 | $0.025 | $0.10, one call |
| Seedream 4.0 | 1 to 4 | $0.0325 | $0.13, one call |
| Flux 2 Pro | 1 to 4 | $0.0375 | $0.15, one call |
Run four takes in parallel
Because each call waits up to 30 seconds in sync mode, a plain loop can take two minutes. Fire the four calls together and read the usage.cost on each response.
import asyncio, os, httpx
async def one(client, prompt):
r = await client.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={"model": "x-ai/grok-image", "prompt": prompt},
)
return r.status_code, r.json()
async def main():
async with httpx.AsyncClient(timeout=60) as client:
out = await asyncio.gather(*[one(client, "a ceramic mug on oak, soft window light") for _ in range(4)])
for status, body in out:
print(status, body.get("usage", {}).get("cost"))
asyncio.run(main())What to check
- A 200 is the image body and a 202 is a job envelope. Read the status code, then poll
/v1/jobs/{id}/statusfor the 202 cases. - Failed or cancelled generations are not billed, so a failed take in the four does not add to the $0.10.
- If you need
nabove 1 on the cheapest tier, Qwen Image has the same $0.025 price with a 4-image ceiling.
Model details and the descriptor format are in the Sume Image API docs.
When one image per call is fine
The cap is not a cost problem, because the per-image price is identical whether you batch or loop. It matters for latency and for your retry logic. A batched n=4 call is one success or failure; four separate calls fail independently, and only the failures are free.
If you are building a picker that shows four takes, run the four calls with a concurrency of four and show each result as it lands. If you are building a bulk pipeline, a work queue with a small worker pool is simpler than one huge gather, and it keeps you below your account's concurrency limits.
Before a large run
Prices and descriptors change when the catalog changes, so confirm them before you spend. Call GET /v1/images/models/{id}/endpoints for the row you plan to use and read its pricing line and supported_parameters; both come back in one response.
Then run a pilot of three to five images and read usage.cost on each response. Multiply by your planned count for a forecast you can trust. Completed generations are billed in full and failed or cancelled ones are not, so a pilot that errors costs nothing.
For big batches, use mode: "async" or mode: "webhook" with a public HTTPS webhook_url, so no request waits on the 30-second sync limit. Poll GET /v1/jobs/{id}/status and fetch GET /v1/jobs/{id}/result when the job completes.
Sources
Related posts
More in Models
- How long can a Seedance 2.5 product clip be on Sume?
On Sume, seedance-2.5 accepts 4 to 30 seconds at 480p, 720p or 1080p. How to pick a length for a product clip and poll the job until it completes.
- Hume Octave 2 covers 11 languages: is yours one?
Hume Octave 2 is reported to cover 11 languages. Before you pick a voice engine, check your language, then verify what Sume's tts_create recorded for the job.
- HunyuanVideo 1.5 on 14 GB of VRAM: run locally or call an API
The HunyuanVideo 1.5 repo lists 8.3B parameters, 480p to 1080p and a 14 GB VRAM minimum with offloading. A local-run versus API checklist.
- HunyuanVideo 1.5 SSTA and FP8: or just call an API
HunyuanVideo-1.5 has 8.3B parameters, SSTA for a 1.87x speedup at 720p, FP8 GEMM and a 14GB VRAM floor. Decide whether to run it or call a hosted video API.
Written by Sume