Image-to-image vs image edit on an API: one route, five fields

On Sume, image-to-image and image edit are one POST /v1/images call. The fields that decide the result: input_references, aspect_ratio, mask_url, quality, n.

5 min readSume
All posts

On Sume, image-to-image and image editing are one operation: you POST to /v1/images with an input_references array, and the model uses your picture as the starting point. What separates a restyle from a precise edit is a handful of fields, not a different endpoint. Five of them decide the result.

One route, two intents

Without input_references, the call generates from text. With them, it edits or restyles. Sume's docs describe the second case as image-to-image, and give a watercolor restyle as the example. The prompt carries the intent: asking for a watercolor look changes everything, asking to swap one object should leave the rest alone.

Reference URLs must be public HTTPS addresses. Sume rejects localhost, private-network and non-HTTPS URLs before it submits the job, so upload your file to storage that has a public link first.

A common mistake is to describe the whole picture in the prompt, as if generating from scratch. For an edit, describe only the change and say what to keep. The model already sees the reference, so repeating its contents spends words and invites drift.

The five fields

The table lists what each field does. The mask and the 16-reference limit are specific to ChatGPT Image 2.5; other models in the catalog list their own input_references ranges, which you read from the catalog before you send.

Read the catalog before you rely on any of these ranges: GET /v1/images/models lists, for each model, which parameters it accepts and the min and max for input_references. A model whose range is 0 to 0 is text-only and rejects references.

Fields that shape an edit on Sume (read 2026-10-07)
FieldWhat it doesNote
input_referencesThe picture or pictures to start fromPublic HTTPS URLs; up to 16 on GPT Image 2.5, up to 10 on many other models
aspect_ratioauto matches the reference's shapeOmitting it is not the same as auto
mask_urlMarks the region to changeGPT Image 2.5 only; public HTTPS URL
qualitylow, medium, high, and on 2.5 also xhigh and maxOmitted means high on 2.5
nHow many resultsGPT Image 2.5 allows up to 4 per call

Choosing a path

Use these rules of thumb, then adjust after one test:

  • Restyle a whole picture: one reference, aspect_ratio auto, no mask, a prompt that names the style.
  • Change one region: add mask_url on ChatGPT Image 2.5 and say what should appear there.
  • Combine looks: pass several references and refer to them in the prompt.
  • Explore: set n to 4 at low quality, then repeat the winner at high.

A request to copy

This curl restyles one photo and keeps its shape. The key comes from the environment, and the body is JSON, since multipart uploads return 415.

curl -sS -X POST https://api.sume.com/v1/images \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-image-2.5",
    "prompt": "Make this scene look like a watercolor painting",
    "aspect_ratio": "auto",
    "quality": "medium",
    "input_references": [
      {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
    ]
  }'

Reading the response

Check the HTTP status. A 200 holds data[].url; a 202 is a job envelope with a status and result address, returned when the work outlasts the 30-second wait. The final amount is in usage.cost, which Sume computes as list price times 1.25, with cent rounding.

If you need transparent output, GPT Image 2.5 takes background set to transparent. Check the result in a viewer on a dark background before you ship it.

Edits that return 202 can be read through the jobs routes, covered in Jobs and results. Keep the job address with your own record of the request.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume