Grok Imagine image edit limit: xAI 5 sources, Sume catalog
xAI's Aug 28 release notes raise image editing to 5 source images. On Sume, Grok is edit-capable with n of 1; read input_references per model.

xAI's release notes say image editing now accepts up to 5 source images per request, up from 3. On Sume, Grok Imagine is the edit-capable x-ai/grok-image row with n capped at 1, and the input_references limit comes from the model's descriptor in GET /v1/images/models, where edit-capable models other than GPT Image 2.5 get a ceiling of 10.
xAI's figure is from its release notes; Sume's from the Image API docs and catalog code, read 2026-09-30. Sume does not document how many sources the Grok provider accepts beyond that descriptor, so test before relying on a high count.
What changed on xAI's side?
The same release notes entry lists five reference images for editing (was 3), plus new aspect ratios 21:9 and 5:2. The notes link it to multi-image editing in xAI's own docs.
What does Sume advertise for Grok?
| Item | xAI release notes | Sume catalog |
|---|---|---|
| Edit sources | Up to 5 (was 3) | Descriptor range 0 to 10 for edit-capable models |
| Images per call | Not in this entry | n maximum 1 |
| Edit support | Yes | Edit-capable row |
How do I find the real limit?
The docs say to read the input_references descriptor per model: text-only models show {"min": 0, "max": 0} and reject references. Too many references returns a 400 naming the maximum. Compare with GPT Image's 16 references.
What if I send a parameter Grok does not list?
It is rejected with 400 unsupported_parameter rather than dropped. Since n is capped at 1, loop your calls for multiple variants; see the n-parameter post.
Sources
Related posts
More in Models
- grok-imagine-image-quality retires Nov 2: what about quality?
xAI retires grok-imagine-image-quality on November 2, 2026 and serves it as grok-imagine-image-2.0 at low quality. Sume's x-ai/grok-image takes no quality.
- Grok Imagine video editing: 8.7 s on xAI, Omni edit on Sume
xAI caps Grok Imagine video edits at 8.7 seconds. Sume's Grok row has no video-to-video; edits go through gemini-omni-flash-1.1 with a video_url.
- Grok Imagine video sound: xAI audio vs Sume's silent row
xAI's Grok Imagine video makes audio unless generate_audio is False. Sume's grok-imagine-video-1.5 row lists audio: false and rejects generate_audio.
- HeyGen ElevenLabs v3 model_id per request vs Sume TTS
HeyGen's speech endpoint takes settings.model_id such as eleven_v3 per request. Sume TTS 1.0 hides model ids; its separate TTS Router lists models explicitly.
Written by Sume