Grok Imagine Image 2.0 policy review: cost of a refused image
xAI says generated media gets content policy review. Sume's docs say failed image generations are not billed; here is what is and is not documented.

If a Grok Imagine Image 2.0 request is refused on content policy grounds, you should not be billed for it through Sume: the Image API docs say a generation is either completed and billed in full, or it fails and is not billed. What the docs do not publish is a Grok-specific error code for a policy refusal, so build your client on the failed status, not on a code you have not seen.
xAI's side is from its image generation guide, read on 2026-10-03, and Sume's from the Image API docs and Errors and credits.
What does xAI say about policy review and training?
The guide states that generated media is subject to content policy review, and that the media is not used for training. It describes grok-imagine-image-2.0 as generating images from text and editing them with natural language, with up to 5 source images in one edit request, and up to 10 images per generation request.
On price, xAI says generation is flat per image regardless of prompt length, while edits bill both the input images and the output images. The page points to xAI's pricing page for the amounts, and we do not repeat a number we did not read there.
What does Sume document for a failed image?
Three rules from the Image API docs apply. Completed generations are billed in full at endpoint pricing. Failed or cancelled generations are not billed, and a failed request returns 502 Bad Gateway. A client that disconnects early is billed as a failed generation, which is to say not at all.
The docs also say Sume does not disclose upstream provider identity, and that model ids follow a catalog. Check GET /v1/images/models for the Grok row you intend to call and read its supported_parameters, because a field the row does not list is rejected with 400 unsupported_parameter.
How should the client treat a refusal?
A policy refusal is a user-correctable condition, not a transient fault. Retrying the same prompt with the same references will fail the same way, and it spends a request and a concurrency slot each time. OpenAI's documentation takes the same position for its own image models, saying user-correctable errors should not be retried automatically.
A sound pattern is to classify the outcome in three buckets and act on each differently.
| Outcome | Documented signal on Sume | What to do |
|---|---|---|
| Finished inside 30 seconds | 200 with data[].url | Store the Sume media URL |
| Still running at 30 seconds | 202 with a job envelope | Poll the status URL or wait for a webhook |
| Job ends failed | Status failed, not billed | Show the reason, change the prompt or inputs, then resubmit with a new key |
| Network timeout on your side | No response | Retry with the same Idempotency-Key and the same body |
Why does the idempotency key matter here?
Two different situations look alike from your code. A refusal needs a changed request and a new Idempotency-Key, because it is a different operation. A timeout needs the same key and the same payload, because it is the same operation. The Sume docs say to reuse a key only for the same operation and payload when retrying after client timeouts.
Mixing those up causes either duplicate spend or a stuck retry loop. Log the key with the job id so support can find a request from either side.
What should you not assume?
Do not assume that a refused edit is cheaper because nothing was drawn. On xAI's side, edits are billed for input and output images, but your Sume bill follows the Sume rule that failed generations are not billed. Do not assume a particular error string either; read the job's failure fields when you hit one and handle the status. And do not assume the grok-imagine-image-quality slug will keep working: xAI's release notes say it is retired on November 2, 2026 and requests then go to grok-imagine-image-2.0 with quality set to low, while the original grok-imagine-image is unaffected.
Sources
Related posts
More in Models
- Higgsfield Soul on Sume: text to image, 1 or 4 images, 720p or 1080p
Soul is a text-to-image row in Sume's image catalog: seven aspect ratios, no references, 1 or 4 images per call, 720p or 1080p. Limits and price.
- Is Ideogram 4.0 open source? Apache code, non-commercial weights
Ideogram 4.0's inference code is Apache 2.0, but the weights on Hugging Face ship under a non-commercial agreement. Selling the images needs a paid licence.
- Irodori-TTS-v4-Large: Japanese cloning, 120 s reference, terms
Irodori-TTS-v4-Large is a 3.29B Japanese TTS model with emoji style control and Gemma terms. What the card says and how Sume's audio tools fit.
- Is FLUX.2 deprecated after FLUX 3? Keep production ids pinned on Sume
BFL's documentation says FLUX.2 remains fully supported for production image work. Pin the model id and read Sume's catalog before changing anything.
Written by Sume