MAI-Image-2.6 output cap is 2,359,296 pixels; Sume sizes differ
MAI-Image-2.6 sets a 2,359,296-pixel ceiling and a 768-pixel minimum edge. Sume sets size per model with tiers, ratios and, for GPT models, custom pixels.
On MAI-Image-2.6 and MAI-Image-2.6-Flash, width times height must not exceed 2,359,296 pixels (1536 by 1536 if square) and each edge must be at least 768; the 2.5 models cap at 1,048,576 (Microsoft Learn, read 2026-10-01). Sume has no single rule: Nano Banana rows use resolution tiers (512, 1K, 2K, 4K; the set varies by model), ChatGPT Image 2.5 takes custom pixels with its own rules, and some models accept only an aspect ratio.
Microsoft's launch post describes the 2.6 pair as offering higher resolutions and dynamic aspect ratios, with support for up to 1.5K resolution (Microsoft AI, read 2026-10-01). If you need a larger print or a 4K hero, that ceiling is the thing to plan around.
How do the size rules compare?
Read them as three different designs: free pixels under a cap, named tiers, or a ratio only.
| Model | Size control | Limit |
|---|---|---|
| MAI-Image-2.6 / Flash | width and height | Each at least 768; product at most 2,359,296; output PNG |
| ChatGPT Image 2.5 on Sume | image_size custom pixels or preset | Edges multiples of 16; max edge 3840; aspect at most 3:1; 655,360 to 8,294,400 pixels |
| Nano Banana 2 / Pro on Sume | resolution tier plus aspect_ratio | Tiers 512, 1K, 2K, 4K (set varies by model) |
| Imagen 4 on Sume | aspect_ratio only | 1:1, 16:9, 9:16, 4:3, 3:4 |
What should I do when the model cap is below my target?
Generate at the largest size the model allows, then upscale. A 1536-pixel square is plenty for a web hero but not for print. On Sume, ChatGPT Image 2.5 can render up to 8,294,400 pixels in one call, and Nano Banana can render 4K tiers, so choose those rows when the first pass has to be large. Otherwise generate smaller and upscale afterwards, as covered in generate at 4K or upscale.
import os
import requests
r = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={
"model": "openai/gpt-image-2.5",
"prompt": "wide studio photo of a ceramic mug on oak, soft window light",
"image_size": "2048x1152",
},
timeout=60,
)
print(r.status_code)
print(r.json())Limits
Large sizes make a request slower and costlier. The docs warn that 4K and high quality are the configurations most likely to pass the 30-second sync wait and return 202 with a job instead of an image, so handle both status codes. A custom size must satisfy the GPT rule exactly or the call is rejected. MAI is not in Sume's image list as of this post; the table compares published limits only.
Sources
Related posts
More in Models
- MAI-Image-2.6 edits take 5 references; Sume takes 10 or 16
MAI-Image-2.6 in Foundry accepts up to five JPEG or PNG reference images per edit. On Sume, input_references tops out at 10, or 16 on GPT Image 2.5.
- MAI-Image-2.6 web_grounding flag: what Sume has instead
MAI-Image-2.6 can pull Bing results into an image when web_grounding is on. Sume has no such flag; here is how to pass current facts in the prompt instead.
- MiniMax H3 camera prompts: lens, movement, exposure wording
fal's H3 prompting guide says H3 reads film vocabulary: lens, rack focus, handheld, grain. Wording that works as a prompt and a Sume request that sends it.
- MiniMax H3 Max Recast API: swap people in a video, fal price vs Sume
H3 Max Recast swaps people in a source video for reference photos, keeping motion, cuts and audio. fal lists $0.30 a second at 768p; what Sume accepts.
Written by Sume