Grok Imagine on Sume: n is 1, so four takes is four calls

Grok Imagine's image row on Sume caps n at 1 and costs $0.025. Four takes means four calls and $0.10. Other models accept n up to 4.

4 min readSume
All posts

Answer

Grok Imagine on Sume (x-ai/grok-image) accepts n from 1 to 1, so one request returns one image and four takes cost four requests at $0.025 each, $0.10 in total. Nearly every other catalog row accepts n up to 4, which batches takes into one call at the same per-image price.

The n ceiling is a capability descriptor, not a guess, and it applies per request. GET /v1/images/models returns n: {type: range, min: 1, max: 1} for Grok Imagine. A request that sets a higher value is rejected rather than trimmed.

n ceiling and price on edit-capable low-cost rows (read 2026-10-04)
Modeln rangePer image4 takes
Grok Imagine1$0.025$0.10, four calls
Qwen Image1 to 4$0.025$0.10, one call
Seedream 4.01 to 4$0.0325$0.13, one call
Flux 2 Pro1 to 4$0.0375$0.15, one call

Run four takes in parallel

Because each call waits up to 30 seconds in sync mode, a plain loop can take two minutes. Fire the four calls together and read the usage.cost on each response.

import asyncio, os, httpx

async def one(client, prompt):
    r = await client.post(
        "https://api.sume.com/v1/images",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
        json={"model": "x-ai/grok-image", "prompt": prompt},
    )
    return r.status_code, r.json()

async def main():
    async with httpx.AsyncClient(timeout=60) as client:
        out = await asyncio.gather(*[one(client, "a ceramic mug on oak, soft window light") for _ in range(4)])
    for status, body in out:
        print(status, body.get("usage", {}).get("cost"))

asyncio.run(main())

What to check

  • A 200 is the image body and a 202 is a job envelope. Read the status code, then poll /v1/jobs/{id}/status for the 202 cases.
  • Failed or cancelled generations are not billed, so a failed take in the four does not add to the $0.10.
  • If you need n above 1 on the cheapest tier, Qwen Image has the same $0.025 price with a 4-image ceiling.

Model details and the descriptor format are in the Sume Image API docs.

When one image per call is fine

The cap is not a cost problem, because the per-image price is identical whether you batch or loop. It matters for latency and for your retry logic. A batched n=4 call is one success or failure; four separate calls fail independently, and only the failures are free.

If you are building a picker that shows four takes, run the four calls with a concurrency of four and show each result as it lands. If you are building a bulk pipeline, a work queue with a small worker pool is simpler than one huge gather, and it keeps you below your account's concurrency limits.

Before a large run

Prices and descriptors change when the catalog changes, so confirm them before you spend. Call GET /v1/images/models/{id}/endpoints for the row you plan to use and read its pricing line and supported_parameters; both come back in one response.

Then run a pilot of three to five images and read usage.cost on each response. Multiply by your planned count for a forecast you can trust. Completed generations are billed in full and failed or cancelled ones are not, so a pilot that errors costs nothing.

For big batches, use mode: "async" or mode: "webhook" with a public HTTPS webhook_url, so no request waits on the 30-second sync limit. Poll GET /v1/jobs/{id}/status and fetch GET /v1/jobs/{id}/result when the job completes.

Sources

Related posts

More in Models

All Models posts

Written by Sume