Grok Imagine multi-image edit: xAI says 5 sources, Sume lists 10
xAI's Imagine docs cap multi-image edits at 5 source images and 10 outputs per request. Sume's Grok Imagine row lists 10 references and one output. Port safely.

If you port an xAI multi-image edit to Sume, send at most 5 reference images and expect one output image per call: xAI's docs cap edits at 5 source images and allow up to 10 outputs per request, while Sume's Grok Imagine row (x-ai/grok-image) lists input_references up to 10 and n of exactly 1. The descriptor numbers match neither side of the xAI limits, so treat 5 as the safe ceiling until you have tested more on your own prompt.
Both facts come from pages read on 2026-10-10: xAI's image generation guide and models page, and Sume's Image API docs plus the catalog code on origin/main.
What the two sides publish
The xAI guide says output count can be configured up to 10 images per request, that multi-image editing supports up to 5 source images, and that image edits are billed for both input and output images. The models page lists grok-imagine-image at $0.02 per image, grok-imagine-image-quality at $0.05 and grok-imagine-image-2.0 at $0.04.
Sume's descriptor for the Grok row says text and image inputs, n from 1 to 1, input_references from 0 to 10, and png, jpeg or webp output. The billed price is $0.025 per image, which is the $0.02 list rate times 1.25. The docs do not say which xAI model id the row calls, so I do not claim that it is any of the three above.
| Limit | xAI docs | Sume x-ai/grok-image |
|---|---|---|
| Outputs per request | up to 10 | 1 |
| Source / reference images | up to 5 for multi-image editing | 0 to 10 listed |
| Price per image | $0.02 / $0.04 / $0.05 by model id | $0.025 billed |
| Input images billed | yes, edits bill input and output | one flat output line |
Port plan
Step one is to read the live descriptor rather than trust this table: GET /v1/images/models/x-ai/grok-image/endpoints shows input_references and n today. Step two is a single test edit with three references. If that works, add references one at a time and keep the largest count that still returns the edit you expect.
Step three is the n limit. If your xAI code asked for four variants in one call, send four requests instead. Sume answers a request outside the range with a 400 that names the field and the allowed range, so the failure is loud rather than a silent clamp.
- Send
aspect_ratio: "auto"only if the row lists it; the Grok row's list starts at 2:1 and has noauto. - Reference URLs must be public HTTPS; Sume rejects localhost, private-network and non-HTTPS URLs before submission.
- Budget one flat $0.025 per output image regardless of how many references you attach.
Why the billing difference matters
xAI bills the input images of an edit as well as the output. Sume's Grok line is a single output_image price in the endpoints route, so a 3-reference edit costs $0.025 per output image on Sume. On xAI the same edit adds input-image charges whose size the page I read does not state, so I cannot give a per-edit total for xAI.
If your edits attach many references and your budget is per output, the flat Sume line is easy to forecast: 200 edits at one output each are 200 x $0.025 = $5.00. Check the endpoints route before you run, because a catalog update can change the line.
What changes in a real edit workflow
A typical xAI edit recipe sends a hero shot plus two or three style or product references and asks for one composite. Moving that to Sume keeps the shape of the call: you put the images in input_references as image_url objects, you pick x-ai/grok-image as the model, and you read data[].url from the response. What you give up is the multi-output convenience, and what you gain is a response that already points at a Sume-hosted copy of the result in data[].url, signed and hosted by Sume.
A 200 response means the image is in the body. If the render runs past 30 seconds the call returns 202 with a job envelope instead, so check the status code before you read the body, and poll GET /v1/jobs/{id}/status as the jobs docs describe. The docs name 4K, high quality and large n as the slow configurations most likely to degrade to a job, so write the client for both outcomes from day one.
Finally, remember the Sume docs line that a model accepts only the values its descriptors list. That sentence is the whole contract: if your port sends a field the Grok row does not advertise, you get a 400 that names it, not a quiet fallback.
Sources
Related posts
More in Models
- Higgsfield Soul on Sume: n must be 1 or 4, n=2 and n=3 return 400
Sume's Higgsfield Soul row accepts n of 1 or 4 only, rejects image_size, and offers 720p or 1080p at $0.005 or $0.0075 per image. Here is how to batch it.
- Can Imagen 4 use my product photo? No: pick an edit-capable row
Imagen 4 Fast and Ultra on Sume are text-to-image only and reject references. Google also says Imagen is shut down in its API. Six rows take a product photo.
- Kling 3.0 10-second clip price on Sume, with other options
A 10-second Kling 3.0 clip bills $2.10 at 1080p and $2.10 at 1080p on Sume. Other models at 1080p for comparison.
- Kling 3.0 12-second clip price on Sume, with other options
A 12-second Kling 3.0 clip bills $2.52 at 1080p and $2.52 at 1080p on Sume. Other models at 1080p for comparison.
Written by Sume