Gemini image multi-turn editing: how to loop edits on Sume

Gemini's image models edit across chat turns. Sume's Image API has no chat session, so you pass the last result URL back as a reference each round.

5 min readSume
All posts

You can keep editing one image across several rounds on Sume, but you do it yourself: send each round as its own call to POST /v1/images and pass the previous result's hosted URL in input_references. There is no chat session to resume.

Google's image generation guide lists "multi-turn conversational iteration" next to text-to-image and text-and-image-to-image editing for its Gemini image models (Nano Banana). That is a feature of Google's chat interface. The Sume Image API request table has no conversation, history or thread field, so the loop below replaces it.

The short version for a team deciding between the two: if your product is a chat interface, you will have to build the memory layer around Sume's endpoint. If your product is a pipeline that makes a fixed series of edits, the stateless shape is easier to test, because every round is one request you can log and replay.

What does Google's multi-turn editing actually give you?

The Google page names three ids: gemini-3.1-flash-lite-image, gemini-3.1-flash-image and gemini-3-pro-image, plus the legacy gemini-2.5-flash-image. It describes conversational iteration as one of the editing modes, and says thinking is on by default and cannot be turned off. The page does not publish a turn limit, so this post makes no claim about one.

The practical benefit of a chat session is that you can say "now make the sofa green" without restating the scene. Without a session, you restate what must stay the same in the prompt and attach the image itself.

How do you replace a chat session on Sume?

Each Sume image response carries data[].url, a Sume-hosted signed URL. The reference rules say input URLs must be public HTTPS, so a hosted result URL can be fed straight back as the next round's reference. See the Image API reference for the exact input_references shape.

Set aspect_ratio to auto on edit rounds. The docs say omitting the field is not the same as auto, and auto is what keeps the reference's shape.

  • Round 1: text-only call, keep data[0].url.
  • Round N: new prompt that states the one change, plus the previous URL as the only reference.
  • Write the preserve list again every round; stateless means the model only sees what you send.
  • Store each round's job id so you can recover from a crash instead of paying twice.

A runnable loop

The sample runs three rounds and prints each URL. It reads the model id from an environment variable so you can pick one from GET /v1/images/models. It only handles the 200 case; a 202 means the job is still running and you should read the result from the job endpoints.

import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODEL = os.environ["SUME_IMAGE_MODEL"]
steps = [
    "A minimalist living room, grey sofa, oak floor, soft daylight",
    "Change only the sofa to deep green. Keep the room, light and floor.",
    "Add one framed print above the sofa. Keep everything else unchanged.",
]
url = None
for prompt in steps:
    body = {"model": MODEL, "prompt": prompt}
    if url:
        body["aspect_ratio"] = "auto"
        body["input_references"] = [{"type": "image_url", "image_url": {"url": url}}]
    r = requests.post("https://api.sume.com/v1/images", headers=H, json=body, timeout=60)
    if r.status_code != 200:
        raise SystemExit(f"{r.status_code}: {r.text[:300]}")
    url = r.json()["data"][0]["url"]
    print(url)

What changes about cost and risk?

Google bills a chat turn as a new generation, and so does Sume: every completed call is billed in full, and a failed call is not billed. A loop of three rounds is three charges. Check the endpoint's pricing lines first, and use one of the cheaper rows for exploratory rounds.

Drift is the real risk. Each round re-encodes the previous output, so small details can slip. The earlier posts on edit drift cover the preserve-list habit; the same applies here.

One more habit helps: keep the prompt for every round in your own log next to the URL it produced. When a client says round three was better than round five, you can branch from the round-three URL without regenerating anything. A chat session makes that kind of branching awkward; a list of URLs makes it trivial.

Also keep an eye on file lifetime. Sume returns hosted URLs rather than inline image data, so the reference you pass in round two is a Sume-hosted file, not something you uploaded. Download the finals you intend to keep, as you would with any generated asset.

Chat editing versus a Sume edit loop (read 2026-10-02)
QuestionGemini chat turnsSume Image API
Where does history live?In the chat sessionIn your code
How is the last image reused?ImplicitlyPrevious data[].url as input_references
Billing unitPer generationPer completed image call
Turn limit published?Not on the page readNot applicable

Where does this fall short?

It is more code than a chat window, and it only works with a model whose input_references descriptor allows at least one reference. Text-only models reject references with a 400. Read the descriptor from the catalog before choosing a model for the loop.

Sources

Related posts

More in Models

All Models posts

Written by Sume