Ideogram 4.5 vs Nano Banana 2 for multi-turn edits: cost per turn
Ideogram 4.5 claims clean multi-turn edits. Compare it with Nano Banana 2 on Sume: price per turn, six-turn totals and a drift test you can run.

Ideogram's launch post for 4.5 on 2026-09-30 claims it is its most precise edit model, with multi-turn edits that do not pile up artifacts. That is a claim about quality, and it is worth testing against the model you use now. Sume lists both Ideogram 4.5 and Nano Banana 2 as edit-capable, so one script can run the same chain through each.
| Model and setting | List | Billed per turn | Six turns |
|---|---|---|---|
| Ideogram 4.5 low | $0.03 | $0.0375 | $0.225 |
| Ideogram 4.5 medium | $0.06 | $0.075 | $0.45 |
| Ideogram 4.5 high | $0.22 | $0.275 | $1.65 |
| Nano Banana 2 (1K) | $0.08 | $0.10 | $0.60 |
Run the same chain through both
Each turn feeds the previous output URL back in as the new source. After six turns, compare the final image with the first on a region that was never meant to change. The script prints the cost and a drift score per model.
import io, os, requests
import numpy as np
from PIL import Image
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
START = "https://example.com/shelf.jpg"
STEPS = ["Make the mug blue.", "Add a plant on the left.", "Make the wall warmer."]
def load(url):
return Image.open(io.BytesIO(requests.get(url).content)).convert("RGB")
for model in ["ideogram/ideogram-v4.5", "google/nano-banana-2"]:
url, cost = START, 0.0
for step in STEPS:
r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=90, json={
"model": model, "prompt": step + " Keep everything else unchanged.",
"input_references": [{"type": "image_url", "image_url": {"url": url}}]})
r.raise_for_status()
url, cost = r.json()["data"][0]["url"], cost + r.json()["usage"]["cost"]
a, b = load(START), load(url).resize(load(START).size)
box = (0, int(a.height * 0.6), a.width, a.height)
d = np.abs(np.asarray(a.crop(box), float) - np.asarray(b.crop(box), float)).mean()
print(model, round(cost, 4), round(d, 2))How to read the result
The bottom 40% is a stand-in for an area your prompts never named; pick the real one for your photos. A lower score means less drift, and a flat score across turns means the chain is stable. Ideogram's claim is its own, and this is the way to check it on your images. The earlier post on four rounds for 30 cents has the Ideogram-only costs.
Sources
Related posts
More in Models
- Kling 3.0 15-second clip: $2.10 silent or $3.15 with sound on Sume
Kling 3 on Sume makes 4 to 15 second clips at 720p or 1080p, with audio priced at $0.21 a second and silence at $0.14. Limits, ratios and how to try it.
- Longest single AI video clip by API in 2026: 30 seconds, two Sume ids
Seedance 2.5 and Wan 3.0 make a 30-second clip in one job on Sume. Kling 3 and MiniMax H3 stop at 15 s and Omni at 10 s. Limits and per-second prices.
- LTX-2.5 audio-to-video: three ways to drive a clip with audio on Sume
LTX-2.5 is not on Sume. To make video that follows a voice or track, use audio references on Seedance 2.5, Wan 3.0 or H3, or H3 Max lip sync. Limits and prices.
- MAI-Voice-2.1 vs Flash: what $7 per million saves on a 60-second short
Flash lists at $15 per million characters and MAI-Voice-2.1 at $22. On a 750-character short that is a fraction of a cent. Math for 1, 30 and 3,000 shorts.
Written by Sume