Grok Imagine multi-image edit: xAI says 5 sources, Sume lists 10

xAI's Imagine docs cap multi-image edits at 5 source images and 10 outputs per request. Sume's Grok Imagine row lists 10 references and one output. Port safely.

5 min readSume
All posts

If you port an xAI multi-image edit to Sume, send at most 5 reference images and expect one output image per call: xAI's docs cap edits at 5 source images and allow up to 10 outputs per request, while Sume's Grok Imagine row (x-ai/grok-image) lists input_references up to 10 and n of exactly 1. The descriptor numbers match neither side of the xAI limits, so treat 5 as the safe ceiling until you have tested more on your own prompt.

Both facts come from pages read on 2026-10-10: xAI's image generation guide and models page, and Sume's Image API docs plus the catalog code on origin/main.

What the two sides publish

The xAI guide says output count can be configured up to 10 images per request, that multi-image editing supports up to 5 source images, and that image edits are billed for both input and output images. The models page lists grok-imagine-image at $0.02 per image, grok-imagine-image-quality at $0.05 and grok-imagine-image-2.0 at $0.04.

Sume's descriptor for the Grok row says text and image inputs, n from 1 to 1, input_references from 0 to 10, and png, jpeg or webp output. The billed price is $0.025 per image, which is the $0.02 list rate times 1.25. The docs do not say which xAI model id the row calls, so I do not claim that it is any of the three above.

Grok Imagine limits, xAI docs vs Sume catalog, read 2026-10-10
LimitxAI docsSume x-ai/grok-image
Outputs per requestup to 101
Source / reference imagesup to 5 for multi-image editing0 to 10 listed
Price per image$0.02 / $0.04 / $0.05 by model id$0.025 billed
Input images billedyes, edits bill input and outputone flat output line

Port plan

Step one is to read the live descriptor rather than trust this table: GET /v1/images/models/x-ai/grok-image/endpoints shows input_references and n today. Step two is a single test edit with three references. If that works, add references one at a time and keep the largest count that still returns the edit you expect.

Step three is the n limit. If your xAI code asked for four variants in one call, send four requests instead. Sume answers a request outside the range with a 400 that names the field and the allowed range, so the failure is loud rather than a silent clamp.

  • Send aspect_ratio: "auto" only if the row lists it; the Grok row's list starts at 2:1 and has no auto.
  • Reference URLs must be public HTTPS; Sume rejects localhost, private-network and non-HTTPS URLs before submission.
  • Budget one flat $0.025 per output image regardless of how many references you attach.

Why the billing difference matters

xAI bills the input images of an edit as well as the output. Sume's Grok line is a single output_image price in the endpoints route, so a 3-reference edit costs $0.025 per output image on Sume. On xAI the same edit adds input-image charges whose size the page I read does not state, so I cannot give a per-edit total for xAI.

If your edits attach many references and your budget is per output, the flat Sume line is easy to forecast: 200 edits at one output each are 200 x $0.025 = $5.00. Check the endpoints route before you run, because a catalog update can change the line.

What changes in a real edit workflow

A typical xAI edit recipe sends a hero shot plus two or three style or product references and asks for one composite. Moving that to Sume keeps the shape of the call: you put the images in input_references as image_url objects, you pick x-ai/grok-image as the model, and you read data[].url from the response. What you give up is the multi-output convenience, and what you gain is a response that already points at a Sume-hosted copy of the result in data[].url, signed and hosted by Sume.

A 200 response means the image is in the body. If the render runs past 30 seconds the call returns 202 with a job envelope instead, so check the status code before you read the body, and poll GET /v1/jobs/{id}/status as the jobs docs describe. The docs name 4K, high quality and large n as the slow configurations most likely to degrade to a job, so write the client for both outcomes from day one.

Finally, remember the Sume docs line that a model accepts only the values its descriptors list. That sentence is the whole contract: if your port sends a field the Grok row does not advertise, you get a 400 that names it, not a quiet fallback.

Sources

Related posts

More in Models

All Models posts

Written by Sume