Grok Imagine's magic wand edit vs Sume's mask_url: who has one
xAI's Image 2.0 edits a selected region and keeps the rest. Sume's explicit mask exists on ChatGPT Image 2.5 only; here is what to send for Grok.

xAI's Imagine Image 2.0 announcement lists a magic wand that edits specific regions while preserving the rest, plus segmentation for selecting precise areas. Those are tools in xAI's product; Sume's API has one explicit region control, mask_url, and its docs limit it to ChatGPT Image 2.5. For Grok on Sume, send the image in input_references and describe the region in words.
xAI's features are from its Image 2.0 announcement, read 2026-10-03. Sume's parameters are from the Image API docs.
What the announcement lists
Beyond the region tools, the page lists background removal that exports subjects on transparent backgrounds, multi-reference editing with up to 5 input images, and smart resize that recomposes an image to any aspect ratio. It says the model launched on August 7, 2026 and is available as grok-imagine-image-2.0 in the API, and it reports a ranking from Arena leaderboards as of the announcement date, which this post does not repeat as a current fact.
| Need | xAI's product | Sume API |
|---|---|---|
| Edit one region | Magic wand, segmentation | mask_url on ChatGPT Image 2.5 only |
| Transparent output | Background removal | background: transparent on ChatGPT Image 2.5 only |
| Several sources | Up to 5 images | input_references, ceiling per model in the catalog |
What to send for Grok on Sume
The docs say mask_url and background are accepted by ChatGPT Image 2.5, and a parameter a model does not list is rejected with 400 unsupported_parameter. So a Grok call carries the source in input_references and a prompt that names the region and what must not change.
import os
import requests
key = os.environ["SUME_API_KEY"]
payload = {
"model": "x-ai/grok-image",
"prompt": "Change only the jacket on the person at left to dark green; leave the rest of the image untouched",
"input_references": [
{
"type": "image_url",
"image_url": {
"url": "https://example.com/photo.png"
}
}
]
}
resp = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {key}"},
json=payload,
timeout=60,
)
if resp.status_code == 200:
print([item["url"] for item in resp.json()["data"]])
elif resp.status_code == 202:
print("still running:", resp.json()["data"]["status_url"])
else:
print(resp.status_code, resp.text)
Limits
A prompt does not guarantee that nothing else moves. If an exact region matters, use openai/gpt-image-2.5 with a mask and compare the unchanged area pixel by pixel. Whether the Grok row accepts the reference count you want is in its input_references descriptor.
Sources
Related posts
More in Comparisons
- grok-voice-transcribe-1.0 ended Oct 2: what a silent reroute means
xAI ended grok-voice-transcribe-1.0 on Oct 2, 2026 and routes it to 2.0 at the same price. How to catch silent model swaps, and what Sume STT 1.0 fixes for you.
- H3 Max Recast vs Higgsfield Genjutsu: 10-second cost on Sume
H3 Max Recast is $3.75 (768p) or $5.625 (1080p) for a 10 second source; Higgsfield Genjutsu motion transfer is $3.975 (480p) or $8.5125 (720p).
- Happy Horse 1.1 at $0.14 and $0.18 a second: nearby Sume rows
fal lists Happy Horse 1.1 at $0.14/s (720p) and $0.18/s (1080p). Sume has no Happy Horse id. Compare list prices to Wan, Omni, H3 Max and Kling.
- Hume Octave overage $0.05 to $0.15 per 1,000 characters vs Sume TTS
Hume's Octave overage runs $0.15 per 1,000 characters on Free down to $0.05 on Business; EVI is $0.04 to $0.07 a minute. Sume TTS is $0.0475 per 1,000.
Written by Sume