Edit an AI image in rounds: feed each result URL back as a reference
Sume image calls are stateless. To refine an image over rounds, send the last result URL as input_references, keep the original too, and track cost.

To refine an image over several rounds on Sume, call POST /v1/images again and send the URL of the previous result as an input_references entry, plus the instruction for this round in prompt. Each call is independent: the Image API has no conversation or response id to chain.
That is simple and cheap to reason about. A round costs the model's price for one image, and a failed or cancelled round is not billed.
The round trip
A completed 200 response carries data[].url, a Sume-hosted signed URL. Reference URLs must be public HTTPS, and a Sume-hosted result URL meets that rule, so you can pass it straight into the next request. If you want to keep a result beyond a single session, download it or save it to your own storage as part of your workflow rather than relying on one signed link.
{
"model": "google/nano-banana-2.1",
"prompt": "Same scene, but change the sofa to deep green velvet. Keep everything else unchanged.",
"aspect_ratio": "auto",
"input_references": [
{ "type": "image_url", "image_url": { "url": "https://media.sume.com/img/EXAMPLE/0.png" } }
]
}Keep the original in the loop
Every round re-renders the picture, so small changes you did not ask for can pile up over several rounds. A practical guard is to send two references from round two on: the original first, the latest result second, and to say in the prompt which one is the source of truth for the parts you want to hold.
That works on every model with at least two reference slots. The catalog ceiling is 10 on most editing models and 16 on ChatGPT Image 2.5. Ideogram 4.5 is different: with references it edits the first image and uses up to 4 more, so put the image you want edited first.
| Model | Base price per image (USD) | Five rounds (USD) | Lists aspect_ratio auto |
|---|---|---|---|
| black-forest-labs/flux.2-pro | 0.0375 | 0.1875 | no |
| bytedance-seed/seedream-4.5 | 0.05 | 0.25 | no |
| google/nano-banana-2.1 | 0.10 | 0.50 | yes |
Watch the aspect ratio
Round after round, the shape should not wander. If the model lists auto, send aspect_ratio: "auto" on each edit so the result matches the reference. If it does not, send the same listed ratio every time. A ratio outside the model's list returns 400 invalid_request, and details.supported shows the accepted values.
When the change is local, such as a sign or a sleeve, consider a masked edit instead. mask_url is available on ChatGPT Image 2.5 only; other models return 400 unsupported_parameter for it.
When a new generation beats another round
Rounds are good for small, specific changes. They are a poor tool when the composition itself is wrong. If round three still has the wrong pose, a new text prompt on a fresh call usually costs the same one image and avoids carrying the old layout forward.
Cost is the other limit. Five rounds on Nano Banana 2.1 cost 0.50 USD at base price, which is the same as five fresh images. Compare the price of a round with the price of simply asking for four new takes with n: 4, and spend the rounds only on images that are nearly right.
A checklist per round
Log the prompt, the reference URLs and the result URL for every round. You can store your own labels in the request metadata object, which Sume keeps on the job and does not send to the provider. With that log, you can go back one round when a change turns out wrong, at the cost of one more image.
- Send the original and the latest result as references.
- Describe one change per round.
- Reuse the same model id and the same ratio setting.
- Stop when two rounds in a row change nothing you care about.
Sources
Related posts
More in Developers
- ElevenLabs made eleven_v4_turbo a default: pin your TTS engine
ElevenLabs set eleven_v4_turbo as a default on Oct 5. Sume carries Sonic engines, not Eleven; here is how to pin an engine id so audio does not drift.
- ElevenLabs dubbing 180-minute app limit vs 3 GB API and Sume detach
ElevenLabs dubs up to 180 minutes in the app or 3 GB via API. Sume detach takes 1,800 s of source and 900 s of output per job, so long videos need ranges.
- ElevenLabs dubbing output: mono v1, stereo v2, and Sume channels
ElevenLabs dubbing v1 returns mono and v2 at most stereo. On Sume, detach keeps source channels or mono, and a concat of mixed layouts fails by design.
- Scribe accepts MP4 and MOV; Sume STT needs an audio URL: detach first
ElevenLabs Scribe takes video files such as MP4 and MOV directly. Sume STT wants a public HTTPS audio_url, so a video goes through audio detach first at $0.01.
Written by Sume