MAI-Image-2.6 edits take 5 references; Sume takes 10 or 16

MAI-Image-2.6 in Foundry accepts up to five JPEG or PNG reference images per edit. On Sume, input_references tops out at 10, or 16 on GPT Image 2.5.

5 min readSume
All posts

Microsoft's MAI-Image-2.6 edits API takes up to five reference images per request, sent as JPEG or PNG files in multipart form data with the image field repeated (Microsoft Learn, read 2026-10-01). On Sume you pass input_references as public HTTPS image URLs, and the ceiling is higher: 10 on most editing models and 16 on ChatGPT Image 2.5 (openai/gpt-image-2.5 and openai/gpt-image-2.5-sunburst). Text-only models accept none.

Sume's image model list does not carry a MAI row today, so this is a comparison of limits, not a way to call MAI through Sume. Microsoft's launch note, Pushing the quality-cost frontier with MAI-Image-2.6, lists multi-reference editing as a feature of both MAI-Image-2.6 and the Flash variant. Check GET /v1/images/models for the live list before you plan around any model.

What are the two limits side by side?

The numbers differ, and so does how you hand the images over. Microsoft wants file uploads; Sume wants URLs it can fetch.

Reference-image limits, read 2026-10-01 from Microsoft Learn and the Sume Image API docs
ItemMAI-Image-2.6 (Foundry preview)Sume Image API
Max references per edit510 default; 16 on GPT Image 2.5; 0 on text-only rows
How images are sentMultipart files, field image repeatedJSON input_references with image_url.url
Accepted formatsJPEG or PNGPublic HTTPS URLs; localhost and private hosts rejected
OutputBase64 PNG in b64_jsonSume-hosted URL in data[].url
Over the limitLearn lists 400 for invalid requestsRejected with a 400 before any billing

How do I read the limit for a Sume model?

Do not hard-code 10 or 16. Each model publishes a range descriptor for input_references, and a request above it is rejected before any generation is billed. This script prints the ceiling for every model your key can see:

import os
import requests

r = requests.get(
    "https://api.sume.com/v1/images/models",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
    refs = m["supported_parameters"].get("input_references", {})
    print(m["id"], refs.get("max", 0))

When does the difference matter?

If your edit uses a product shot, a logo, a model photo, a background and a style board, that is five inputs and fits either side. A catalog collage with eight or twelve source photos only fits a Sume model with a higher ceiling, which means GPT Image 2.5. A single model with five is fine for most brand-lock edits.

Whichever limit you hit, order matters less than clarity: say in the prompt what each reference is for (image 1 is the product, image 2 is the room). See GPT Image 2.5 with 16 references for a prompt layout.

Limits of this comparison

Microsoft's model is a public preview with no SLA, per the Learn page, and its limits can change. Sume's ceilings come from the catalog descriptors and can change as rows are added. Neither number says anything about quality: more references does not mean a better edit, and a model can ignore references it finds redundant. Test with your own inputs, and read the Image API reference for reference URL rules.

Sources

Related posts

More in Models

All Models posts

Written by Sume