Which Sume image model for a photo edit: mask, references or pixels?
Pick between ChatGPT Image 2.5, Ideogram 4.5 and Nano Banana 2 for a photo edit on Sume by what the edit needs: a mask, many references, or untouched pixels.
Start from what the edit needs
Three needs decide the model for a photo edit: a mask over one area, a pile of reference images, or a promise that untouched pixels stay as they were. Sume's catalog covers all three, and a short script reads the catalog so you do not have to remember which model has which field.
Sume's Image API docs say a model only accepts the parameters its catalog descriptors list, so read supported_parameters before sending a field. A field the model does not list returns 400 unsupported_parameter.
Vendor pages say what each model is built for. Google's Gemini API docs list Nano Banana 2 as gemini-3.1-flash-image with up to 10 object images and up to 4 character images. Ideogram's API overview says pixels an edit does not touch are copied exactly and allows up to four reference images. OpenAI's guide covers masks for GPT Image.
Ask the catalog which models take a mask
GET /v1/images/models returns each model with a supported_parameters object. This script lists the ids that publish mask_url; change the field name to background or input_references to answer the other two questions. Per the docs, mask_url and background are live on ChatGPT Image 2.5 only.
import os
import requests
resp = requests.get(
"https://api.sume.com/v1/images/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
resp.raise_for_status()
for model in resp.json()["data"]:
if "mask_url" in model["supported_parameters"]:
print(model["id"])A decision table from the vendor pages and Sume docs
Use this as a starting point and run the same photo on two models for anything important.
| Need | Model on Sume | Why |
|---|---|---|
| Mask over one region | openai/gpt-image-2.5-sunburst | mask_url is live on ChatGPT Image 2.5 only |
| Many references | openai/gpt-image-2.5-sunburst | Up to 16 references on this model |
| Exact untouched pixels | ideogram/ideogram-v4.5 | Vendor says untouched pixels are copied exactly |
| Object and character sets | google/nano-banana-2 | Google lists 10 object and 4 character images |
Test with your own photo
Tables only narrow the choice. Run your actual photo on two models with the same prompt and compare the areas you care about, such as text, edges and faces. Failed generations are not billed, so a rejected request costs nothing.
Quality and the first try
On ChatGPT Image 2.5 the quality field takes auto, low, medium, high, xhigh or max, and leaving it out means high. For a first pass at a layout idea, a lower tier is a reasonable way to look at composition before you pay for a final render.
Keep the source photo, the prompt and the response together for each option. That makes it easy to rerun the one you pick at a higher quality tier.
Sync, jobs and the bill
Treat the response code as the switch. 200 means the image body is in the response. 202 means a job was created because the 30-second wait ran out, and the image is read later from GET /v1/jobs/{id}/result.
A completed image is billed in full and a failed or cancelled one is not. The charge shown in usage.cost is provider list price times 1.25, so you can log it per edit and sum a batch from those numbers.
Sources
Related posts
More in Models
- Which Sume image models accept quality? Only five do
Only five Sume image catalog rows list a quality field: GPT Image 2, 2.5 and Sunburst, Ideogram V3 and 4.5. The rest return 400 unsupported_parameter.
- Which Sume image models can't edit? Five text-to-image-only rows
Five Sume image rows take text prompts only and reject image_urls: Soul, Imagen 4 Fast and Ultra, Recraft V4 and Qwen Image Max. Prices and what to use instead.
- Which Sume image models make 2K or 4K output, by model
FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.
- Which Sume video model fits your inputs: text, photo, clip, audio
Match the input you hold to a Sume video model: prompt, first frame, end frame, references, audio sample, or a clip to edit. With the 400s each mix causes.
Written by Sume