MAI-Image-2.6 edits take 5 references; Sume takes 10 or 16
MAI-Image-2.6 in Foundry accepts up to five JPEG or PNG reference images per edit. On Sume, input_references tops out at 10, or 16 on GPT Image 2.5.

Microsoft's MAI-Image-2.6 edits API takes up to five reference images per request, sent as JPEG or PNG files in multipart form data with the image field repeated (Microsoft Learn, read 2026-10-01). On Sume you pass input_references as public HTTPS image URLs, and the ceiling is higher: 10 on most editing models and 16 on ChatGPT Image 2.5 (openai/gpt-image-2.5 and openai/gpt-image-2.5-sunburst). Text-only models accept none.
Sume's image model list does not carry a MAI row today, so this is a comparison of limits, not a way to call MAI through Sume. Microsoft's launch note, Pushing the quality-cost frontier with MAI-Image-2.6, lists multi-reference editing as a feature of both MAI-Image-2.6 and the Flash variant. Check GET /v1/images/models for the live list before you plan around any model.
What are the two limits side by side?
The numbers differ, and so does how you hand the images over. Microsoft wants file uploads; Sume wants URLs it can fetch.
| Item | MAI-Image-2.6 (Foundry preview) | Sume Image API |
|---|---|---|
| Max references per edit | 5 | 10 default; 16 on GPT Image 2.5; 0 on text-only rows |
| How images are sent | Multipart files, field image repeated | JSON input_references with image_url.url |
| Accepted formats | JPEG or PNG | Public HTTPS URLs; localhost and private hosts rejected |
| Output | Base64 PNG in b64_json | Sume-hosted URL in data[].url |
| Over the limit | Learn lists 400 for invalid requests | Rejected with a 400 before any billing |
How do I read the limit for a Sume model?
Do not hard-code 10 or 16. Each model publishes a range descriptor for input_references, and a request above it is rejected before any generation is billed. This script prints the ceiling for every model your key can see:
import os
import requests
r = requests.get(
"https://api.sume.com/v1/images/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
refs = m["supported_parameters"].get("input_references", {})
print(m["id"], refs.get("max", 0))When does the difference matter?
If your edit uses a product shot, a logo, a model photo, a background and a style board, that is five inputs and fits either side. A catalog collage with eight or twelve source photos only fits a Sume model with a higher ceiling, which means GPT Image 2.5. A single model with five is fine for most brand-lock edits.
Whichever limit you hit, order matters less than clarity: say in the prompt what each reference is for (image 1 is the product, image 2 is the room). See GPT Image 2.5 with 16 references for a prompt layout.
Limits of this comparison
Microsoft's model is a public preview with no SLA, per the Learn page, and its limits can change. Sume's ceilings come from the catalog descriptors and can change as rows are added. Neither number says anything about quality: more references does not mean a better edit, and a model can ignore references it finds redundant. Test with your own inputs, and read the Image API reference for reference URL rules.
Sources
Related posts
More in Models
- MAI-Image-2.6 web_grounding flag: what Sume has instead
MAI-Image-2.6 can pull Bing results into an image when web_grounding is on. Sume has no such flag; here is how to pass current facts in the prompt instead.
- MiniMax H3 camera prompts: lens, movement, exposure wording
fal's H3 prompting guide says H3 reads film vocabulary: lens, rack focus, handheld, grain. Wording that works as a prompt and a Sume request that sends it.
- MiniMax H3 Max Recast API: swap people in a video, fal price vs Sume
H3 Max Recast swaps people in a source video for reference photos, keeping motion, cuts and audio. fal lists $0.30 a second at 768p; what Sume accepts.
- MiniMax H3 sound design prompts: direct the audio like the picture
fal's H3 guide says to direct audio as deliberately as picture: name sonic elements, not 'music'. What it looks like in a Sume request, and what you can't set.
Written by Sume