input_references by Sume image model: 0, 5, 10 or 16 images
Sume image rows take 0, 5, 10 or 16 reference images. A table for every row, plus how Nano Banana 2.1's 14 and Hy Image 3.5's 20 compare.

Sume's image rows accept 0, 5, 10 or 16 reference images in input_references, depending on the row. ChatGPT Image 2.5 takes 16, Ideogram 4.5 takes 5 (the first is the image to edit and up to 4 more are references), most edit-capable rows take 10, and five rows take none. The numbers come from the Sume Image API docs and the repository catalog on 2026-10-08.
The comparison got more interesting this week. Fal's page for the Nano Banana 2.1 edit endpoint lists up to 14 reference images, and OpenRouter's listing for Tencent's Hy Image 3.5 Preview says up to 20. Both exceed the Sume cap for the same model families, or are not on Sume at all, so plan the job around the cap, not the vendor number.
The cap on every row
| Cap | Rows |
|---|---|
| 16 | openai/gpt-image-2.5 and openai/gpt-image-2.5-sunburst |
| 10 | google/nano-banana-2.1, google/nano-banana-pro, openai/gpt-image-2, Seedream 5.0 Lite, 4.5 and 4.0, FLUX.2 pro and flex, Qwen Image, Grok Image, Ideogram V3 |
| 5 | ideogram/ideogram-v4.5 (1 source image plus up to 4 references) |
| 0 (text only) | higgsfield/soul, Imagen 4 Fast and Ultra, Recraft V4, Qwen Image Max |
How to send them
Each reference is an object, and the URL must be public HTTPS. Sume rejects localhost, private-network and non-HTTPS URLs before it submits the job, and an unreachable URL fails with input_media_unreachable.
{
"model": "google/nano-banana-2.1",
"prompt": "Put the mug from image 1 on the desk from image 2",
"aspect_ratio": "auto",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/mug.png"}},
{"type": "image_url", "image_url": {"url": "https://example.com/desk.jpg"}}
]
}When the vendor number is bigger
- Nano Banana 2.1 edit: 14 at Fal, 10 on Sume. Send your 10 most important images and describe the rest in the prompt.
- Hy Image 3.5 Preview: 20 on OpenRouter's listing. It is not in the Sume catalog, so the closest high-count row is ChatGPT Image 2.5 at 16.
- Ideogram 4.5: five total, because one slot is the image you edit. A job with a logo and a style board fits; a ten-image moodboard does not.
Cost and shape
References do not change the per-image price on most rows, but GPT Image 2.5 admission includes estimated input tokens, so a 16-image job reserves more than a text-only one. On edits, set aspect_ratio: "auto" to keep the shape of the source; if you omit it the result is not the same as auto.
A mask is a separate matter. mask_url works only on ChatGPT Image 2.5. For other rows you describe the region in words.
Planning an edit around the cap
Count your inputs before you write the prompt. A product shot usually needs the product, a background and a style example: three images. A character sheet might use six. A brand kit with a logo, two fonts shown as images and four photos hits seven, which fits ten-reference rows but not Ideogram 4.5.
Refer to the images in the prompt by order, for example 'image 1' and 'image 2', and keep that order stable across a batch. If a row rejects the request because of the cap, the error is a 400 that names the field, so you can catch it in a test before a long run. The descriptor is a range with min 0 and a max, which you can read from GET /v1/images/models instead of hard-coding the table above.
What Sume does not do
Sume does not merge, reorder or pre-crop references, and it does not lift its cap because a vendor allows more. A request over the cap is rejected by the catalog check rather than trimmed.
Sources
Related posts
More in Comparisons
- Is 1080p worth it for 30 s AI shots? 2.46x on Seedance, 2x on Wan
On Sume, 1080p costs 2.46 times 720p for a 30 s Seedance 2.5 clip and exactly 2 times on Wan 3.0. A rule for choosing the resolution.
- Is MAI-Voice-2.1 on Sume? What the TTS Router lists instead
No. Sume's TTS Router lists five Sonic ids, each up to 20,000 characters at the same price. MAI-Voice-2.1 and Flash are not in the catalog.
- Is Wan 3.0 an upgrade on Wan 2.7? 113 Elo more on the AA editing board
On the AA video-editing board Wan 3.0 (1,187, $12.00 a minute) beats Wan 2.7 (1,074, $16.90) by 113 Elo. Sume lists Wan 3.0, not 2.7; it runs 2 to 30 s.
- JSON2Video credits per second and 4K at 4x vs Sume Timeline
JSON2Video deducts one credit per video second and four times that for 4K. Sume Timeline charges $0.10 per rounded output minute. What a minute costs on each.
Written by Sume