Gemini image multi-turn editing: how to loop edits on Sume
Gemini's image models edit across chat turns. Sume's Image API has no chat session, so you pass the last result URL back as a reference each round.

You can keep editing one image across several rounds on Sume, but you do it yourself: send each round as its own call to POST /v1/images and pass the previous result's hosted URL in input_references. There is no chat session to resume.
Google's image generation guide lists "multi-turn conversational iteration" next to text-to-image and text-and-image-to-image editing for its Gemini image models (Nano Banana). That is a feature of Google's chat interface. The Sume Image API request table has no conversation, history or thread field, so the loop below replaces it.
The short version for a team deciding between the two: if your product is a chat interface, you will have to build the memory layer around Sume's endpoint. If your product is a pipeline that makes a fixed series of edits, the stateless shape is easier to test, because every round is one request you can log and replay.
What does Google's multi-turn editing actually give you?
The Google page names three ids: gemini-3.1-flash-lite-image, gemini-3.1-flash-image and gemini-3-pro-image, plus the legacy gemini-2.5-flash-image. It describes conversational iteration as one of the editing modes, and says thinking is on by default and cannot be turned off. The page does not publish a turn limit, so this post makes no claim about one.
The practical benefit of a chat session is that you can say "now make the sofa green" without restating the scene. Without a session, you restate what must stay the same in the prompt and attach the image itself.
How do you replace a chat session on Sume?
Each Sume image response carries data[].url, a Sume-hosted signed URL. The reference rules say input URLs must be public HTTPS, so a hosted result URL can be fed straight back as the next round's reference. See the Image API reference for the exact input_references shape.
Set aspect_ratio to auto on edit rounds. The docs say omitting the field is not the same as auto, and auto is what keeps the reference's shape.
- Round 1: text-only call, keep
data[0].url. - Round N: new prompt that states the one change, plus the previous URL as the only reference.
- Write the preserve list again every round; stateless means the model only sees what you send.
- Store each round's job id so you can recover from a crash instead of paying twice.
A runnable loop
The sample runs three rounds and prints each URL. It reads the model id from an environment variable so you can pick one from GET /v1/images/models. It only handles the 200 case; a 202 means the job is still running and you should read the result from the job endpoints.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODEL = os.environ["SUME_IMAGE_MODEL"]
steps = [
"A minimalist living room, grey sofa, oak floor, soft daylight",
"Change only the sofa to deep green. Keep the room, light and floor.",
"Add one framed print above the sofa. Keep everything else unchanged.",
]
url = None
for prompt in steps:
body = {"model": MODEL, "prompt": prompt}
if url:
body["aspect_ratio"] = "auto"
body["input_references"] = [{"type": "image_url", "image_url": {"url": url}}]
r = requests.post("https://api.sume.com/v1/images", headers=H, json=body, timeout=60)
if r.status_code != 200:
raise SystemExit(f"{r.status_code}: {r.text[:300]}")
url = r.json()["data"][0]["url"]
print(url)What changes about cost and risk?
Google bills a chat turn as a new generation, and so does Sume: every completed call is billed in full, and a failed call is not billed. A loop of three rounds is three charges. Check the endpoint's pricing lines first, and use one of the cheaper rows for exploratory rounds.
Drift is the real risk. Each round re-encodes the previous output, so small details can slip. The earlier posts on edit drift cover the preserve-list habit; the same applies here.
One more habit helps: keep the prompt for every round in your own log next to the URL it produced. When a client says round three was better than round five, you can branch from the round-three URL without regenerating anything. A chat session makes that kind of branching awkward; a list of URLs makes it trivial.
Also keep an eye on file lifetime. Sume returns hosted URLs rather than inline image data, so the reference you pass in round two is a Sume-hosted file, not something you uploaded. Download the finals you intend to keep, as you would with any generated asset.
| Question | Gemini chat turns | Sume Image API |
|---|---|---|
| Where does history live? | In the chat session | In your code |
| How is the last image reused? | Implicitly | Previous data[].url as input_references |
| Billing unit | Per generation | Per completed image call |
| Turn limit published? | Not on the page read | Not applicable |
Where does this fall short?
It is more code than a chat window, and it only works with a model whose input_references descriptor allows at least one reference. Text-only models reject references with a 400. Read the descriptor from the catalog before choosing a model for the loop.
Sources
Related posts
More in Models
- gemini-omni-1.1-flash vs gemini-omni-flash-1.1: which id goes where
Google writes the model id gemini-omni-1.1-flash; Sume's catalog id is gemini-omni-flash-1.1. A mapping table, the old preview id, and how to avoid a typo.
- Gemini Omni audio reference: unsupported; Sume models that take audio
Google says Omni's API doesn't accept uploaded audio references. Sume's Omni row has no reference_audio_urls; Seedance 2.x, Wan 3.0 and MiniMax H3 honor audio.
- Gemini Omni extend: no new dialogue on an uploaded talking clip
Google says you can't extend an uploaded Omni clip where someone talks to add dialogue; extension only appends, to clips up to 10 s. Sume lists no extend mode.
- Gemini Omni with several videos: 3 references, no cross-video use
Gemini Omni takes up to 3 reference clips of 3 s each, yet Google warns that reasoning across several videos may degrade output. Sume: one video_url source.
Written by Sume