GPT Image 2.5 comments on an image: review notes as mask edits
Sume's image API has no comment field. To act on several notes about one picture, send one masked edit per note and feed each result into the next request.

Sume's image API has no comment or annotation field: POST /v1/images takes a prompt, reference images and an optional mask_url. To act on several review notes about one picture, turn each note into its own masked edit. Send the picture as input_references, the region as mask_url, the note as the prompt, and pass each result into the next request.
The fields come from Sume's Image API docs, read 2026-09-29.
mask_url is listed on the two GPT Image 2.5 ids, openai/gpt-image-2.5 and openai/gpt-image-2.5-sunburst.
How does a note map to request fields?
A note has two parts, where and what. The where becomes a mask image you host at a public HTTPS URL; the what becomes the prompt. Keep one region per note so a mask never covers two different fixes.
| Part of the note | Request field |
|---|---|
| The picture being reviewed | input_references |
| The marked region | mask_url (public HTTPS) |
| The requested change | prompt |
| Output shape follows the source | aspect_ratio: "auto" |
How do I apply several notes in order?
Loop over the notes. Each response returns a Sume-hosted URL in data[].url, and that URL becomes the reference for the next note, so later edits build on earlier ones. Reference and mask URLs must be public HTTPS, so host the masks before you start.
MASKS=(https://example.com/mask-logo.png https://example.com/mask-sky.png)
NOTES=("Redraw the logo crisply, same colors" "Change the sky to dusk")
cur=https://example.com/draft.png
for i in 0 1; do
cur=$(curl -s -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\":\"openai/gpt-image-2.5\",\"prompt\":\"${NOTES[$i]}\",\"mask_url\":\"${MASKS[$i]}\",\"aspect_ratio\":\"auto\",\"input_references\":[{\"type\":\"image_url\",\"image_url\":{\"url\":\"$cur\"}}]}" \
| jq -r '.data[0].url')
echo "$i: $cur"
doneWhich mask field name is right?
On /v1/images the field is mask_url. The older Image 1.0 page names its mask field mask_image_url, and a parameter the selected model does not list is rejected with 400 unsupported_parameter rather than silently dropped. Check the spelling against the endpoint you call.
What if one edit drifts outside its mask?
Compare each output with its input before you feed it forward, and retry the single note that drifted instead of restarting the chain. Blocking calls wait 30 seconds at most, so a slow edit can return a 202 job; poll it and continue once its result is ready.
How do I make the mask and the note agree?
Draw the mask around the smallest area the note is about, and word the prompt about that area only. A note such as "make the sky dusk" pairs with a sky mask; a prompt that also mentions the logo invites changes elsewhere. Say what must stay the same, for example "keep the buildings unchanged", so the instruction is explicit.
Host each mask at its own public HTTPS URL and name the files after the notes, so a failed step is easy to trace. If two notes overlap the same region, merge them into one note instead of running them as separate edits, since the second edit would repaint what the first just fixed.
Sources
Related posts
More in Use cases
- Sketch to image with GPT Image 2.5 API: send a drawing as a reference
To turn a sketch into a finished image with GPT Image 2.5 on Sume, send the drawing in input_references and describe the result. The request, size tips, limits.
- Gym promo video with AI: from your own gym photos
Make a gym promo video without a shoot: animate photos of your own floor and classes, add a voiced offer, music and captions, and render a vertical cut.
- Hair salon promo video with AI, from your own photos
Make a hair salon promo video from your own photos: animate the space and real finished looks, add a voiced offer and music, and render it vertical.
- Hotel promotional video with AI, from your own photos
A hotel promotional video can be made from the photos you already have: animate each shot, add a voiced welcome and music, render wide and vertical cuts.
Written by Sume