Hy Image 3.5 Preview 20 references vs GPT Image 2.5 16: which edits?
Hy Image 3.5 Preview lists up to 20 references; ChatGPT Image 2.5 on Sume takes 16, plus a mask_url. A comparison table for multi-image edits, read 2026-10-07.

For a multi-image edit today, ChatGPT Image 2.5 is the one you can call on Sume, with up to 16 references and an optional mask_url; Tencent Hy Image 3.5 Preview lists up to 20 references on its OpenRouter page but is not in Sume's catalog. If your brief needs more than 16 inputs, Hy is the only one of the two that lists the room, and you would call it through its own listing.
The difference of four references only matters for sheets such as a lookbook or a character board. For a normal product edit with one photo and one logo, both are far above what you need.
Side by side
The Sume column comes from the Image API docs. The Hy column comes from the OpenRouter listing read today. Cells that a source does not state are marked as not stated, rather than filled in.
| Feature | ChatGPT Image 2.5 on Sume | Hy Image 3.5 Preview (OpenRouter) |
|---|---|---|
| References | Up to 16 | Up to 20 |
| Mask input | Optional mask_url (public HTTPS) | Not stated on the listing |
| Multi-turn editing | Edit loop by chaining calls | Listed as a mode |
| Output size | Custom edges up to 3840, multiples of 16 | Up to 4K |
| Price basis | Image tokens: $30 per million output, $8 per million input, before Sume pricing | $1.60 per million tokens |
| Callable on Sume | Yes: openai/gpt-image-2.5 | No, not in the catalog |
What a reference actually costs on Sume
On ChatGPT Image 2.5, references are input image tokens at $8 per million before Sume pricing, per the Fal token rates the docs cite. Input token counts are estimates, so each added reference raises the estimate and the more you send the more the call costs. Send only the images that change the result.
- Put the image to edit first and style or identity references after it.
- Use
mask_urlwhen only one region should change. - Keep reference URLs public HTTPS; Sume rejects private, localhost and non-HTTPS URLs before submission.
If you have 17 to 20 inputs
More references only help if the model uses them. On Sume the ceiling is 16 per call, so a 17-to-20 image brief has to be reduced or split into more than one call. Price both options before you choose, since each call is a separate generation.
- Rank references by how much each changes the result.
- Drop near-duplicates before counting.
- Compare the result with a single call on one real brief.
How to choose
If the edit is a swap inside one frame, choose the model with a mask and a price you can read in the catalog. If it is a collage of many inputs where a 17th reference matters, test Hy directly and measure its token use, because the listing gives no per-image figure.
For a ChatGPT Image 2.5 reference cost grid with the input token math, read the 16 references versus one reference post.
The reference count is a ceiling, not a target. Run the same brief on both models with the same few references, compare the results, and add the cost of failed attempts before you decide.
Sources
Related posts
More in Comparisons
- Ideogram 4.5 edits: default medium, 5 references, vs GPT Image 2.5
Ideogram 4.5 on Sume edits the first of up to 5 references and defaults to medium quality; ChatGPT Image 2.5 takes 16 and defaults to high. Side by side.
- Image model bake-off: 20 prompts on 6 Sume models costs $6.125
A fair test sends the same 20 prompts to each model. On Sume, Grok, Qwen, Flux 2 Pro, Seedream 5 Lite, Ideogram 4.5 and Nano Banana 2.1 cost $6.125 in all.
- Instagram's Reels page says 3 and 20 minutes: which one to plan for
Instagram's Reels features page says both multi-clip videos up to 3 minutes and clips adding up to 20 minutes. Plan creative for 3, tools for 20, and read both.
- Instagram Series or YouTube Shorts series: render one master
Instagram's Series test and YouTube Shorts series both want vertical video. Render one master, trim it per platform, and keep an episode sheet. Rates included.
Written by Sume