GPT Image 2.5 edits through Sume: mask_url and 16 references
Sume's GPT Image 2.5 takes up to 16 references, an optional mask_url and background auto, transparent or opaque. OpenAI: transparent needs png or webp.

For GPT Image 2.5 edits, Sume exposes three controls in one request: up to 16 input_references, an optional mask_url, and background set to auto, transparent or opaque. OpenAI's guide says a transparent background needs png or webp output, so set output_format to one of those when you ask for transparency. The limits below are from the two pages. OpenAI's page says up to four input images where Sume documents 16, so test your own reference count.
What each side documents
| Control | OpenAI guide | Sume |
|---|---|---|
| Reference images | Page text says up to four input images | Up to 16 input_references, public HTTPS URLs |
| Mask | Masks are supported for edits | Optional mask_url, public HTTPS, GPT Image 2.5 edits |
| Background | Transparent works with png or webp | auto, transparent, opaque |
| Quality | low, medium, high, xhigh, max, auto | Same set; default high; auto reserves max |
| Size | Multiples of 16, ratio 1:3 to 3:1, 655,360-8,294,400 pixels | Same multiple-of-16 and ratio rules, max edge 3840 |
| Output format | png, jpeg, webp; transparency needs png or webp | png, jpeg, webp |
An edit request
Sume rejects localhost, private-network and non-HTTPS reference URLs before submission. If a model's input_references range is 0 to 0, it is text-to-image only. Also set aspect_ratio: "auto" on edits to match the reference. Omitting it is not the same as auto.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: mug-edit-001" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Replace the label with a plain white label",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/mug.jpg"}}
],
"mask_url": "https://example.com/mug-mask.png",
"aspect_ratio": "auto",
"background": "transparent",
"output_format": "png",
"mode": "async"
}'Cost and timing
Flare and Sunburst share the same token rates in the Sume docs: $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens, before Sume pricing. The docs give 1024x1024 output estimates of $0.09366 at xhigh and $0.21072 at max, before input tokens. Check the per-request estimate with the catalog pricing for your size and quality rather than computing it yourself.
Edits with high quality can run long. See GPT Image 2.5 can take 2 minutes: submit async for the async pattern.
Limits to remember
mask_urlis documented for GPT Image 2.5 only. Other models reject parameters their catalog descriptors do not list.- Sume's transparent-still guidance for older paths points to Image 1.0 with
transparency: true. - The Sume docs do not describe how the mask is interpreted. Test with a known mask before you build a pipeline on it.
Planning the inputs
Count your references before you submit. Sume caps GPT Image 2.5 at 16, so a catalog of 20 product angles needs a selection step on your side. Host every reference and the mask at a public HTTPS URL, because localhost, private-network and non-HTTPS addresses are rejected before submission.
For size, both pages use multiples of 16 and a 3:1 ratio limit. A 2,048 by 1,152 image passes: 2,048 / 16 = 128 and 1,152 / 16 = 72, and its pixel count of 2,359,296 sits inside Sume's 655,360 to 8,294,400 range. Sume's 3,840 pixel maximum edge applies to both sides. Check any custom size against these rules before you pay for a request that will be rejected.
Sources
Related posts
More in Models
- Pocket TTS license: the repo says MIT, not Apache-2.0
Some roundups call Kyutai's Pocket TTS Apache-2.0. Its GitHub page and LICENSE file read MIT-style. How to check a TTS license before you build on it.
- Pocket TTS runs on 2 CPU cores: what ~200 ms first audio means
Kyutai's Pocket TTS lists 100M parameters, 2 CPU cores and ~200 ms to first audio. Whether that matters for a video voiceover, and a hosted TTS job's numbers.
- Qwen-Image 2.1 native RGBA vs Sume's ChatGPT Image 2.5 transparency
Qwen-Image 2.1's card lists native RGBA transparency under a research licence. On Sume, transparent output comes from ChatGPT Image 2.5's background field.
- Qwen-Image 2.1 is 7B: self-host it or call hosted qwen-image on Sume
Qwen-Image 2.1 is a 7B model under the Qwen Research License. Sume hosts qwen-image and qwen-image-max, not 2.1; hosted costs $0.025 or $0.094 per image.
Written by Sume