Nano Banana edit slow? Thinking can't be turned off, expect 202
Google says Nano Banana thinking cannot be disabled. On Sume a call that outlasts 30 seconds returns 202 with a job id. How to handle both without paying twice.

If a Nano Banana call is slower than you expected, part of the reason is on Google's side: thinking is on by default and cannot be disabled. On Sume that shows up as a 202 job envelope once a call passes the 30-second blocking budget, and the fix is to read the job, not to resend the request.
Google's image generation guide says the thinking process is "enabled by default; cannot be disabled," that the model "generates interim test images before final output," and that levels are minimal (the default) and high. The Sume Image API parameter table has no thinking field, so you cannot select high through it.
What does the Sume side look like?
POST /v1/images blocks for up to 30 seconds. If the generation is done, you get 200 with data[].url. If not, you get 202 with a job id, a status_url and a result_url. The docs name 4K, high quality and large n as the configurations most likely to degrade to 202.
So the status code, not the body shape, tells you what you have. A 202 is a normal answer, not an error.
How do you handle both cases?
Submit once, branch on the code, poll the status with backoff, and never resubmit a paid request because your own process timed out. The Jobs docs say exactly that. If you do not want to poll, send mode: "webhook" with a public HTTPS webhook_url; see the webhooks guide.
If you are building an interactive tool, show the user a progress state as soon as the 202 arrives, then swap in the picture when the status flips. Users tolerate a 40-second generation far better than a spinner that ends in a timeout and a second charge.
import os, time, requests
B = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def find_status(o):
if isinstance(o, dict):
if isinstance(o.get("status"), str): return o["status"]
for v in o.values():
s = find_status(v)
if s: return s
r = requests.post(B + "/v1/images", headers=H, timeout=60, json={
"model": os.environ["SUME_IMAGE_MODEL"], "resolution": "4K",
"prompt": "Studio photo of a ceramic mug on a marble counter"})
if r.status_code == 202:
d = r.json()["data"]
delay = 2
while find_status(requests.get(d["status_url"], headers=H).json()) not in (
"completed", "failed", "canceled"):
time.sleep(delay); delay = min(delay * 2, 20)
print(requests.get(d["result_url"], headers=H).json())
else:
print(r.status_code, r.json())What should you not do?
Do not retry the POST when you see 202; the job is already accepted and may be billed once it completes. Failed generations are not billed, so a failed status is safe to retry, ideally with a new prompt or a cheaper tier.
- Prefer 2K over 4K for drafts to stay under the blocking budget.
- Lower
nwhen you only need one candidate. - Use webhook mode for batch jobs so no worker sits waiting.
- Check the model's
resolutiondescriptor first; a lite model may offer only 1K.
Is slow always thinking?
No. Queue time, provider load and size all add to it, and the job events timeline shows where the time went. Google's page explains one cause for the Gemini family, not every delay.
The Google page also says Lite is limited to 1K and that thinking generates interim test images, which is a plausible reason a quick prompt can still take a while. Treat that as context, not a measurement; this post does not publish timings because none were measured.
Sources
Related posts
More in Developers
- Netflix subtitle limit: 42 characters per line, and max_chars
Netflix's English timed text spec allows 42 characters per line and two lines. Sume's caption design.phrasing.max_chars accepts 4 to 60, so 42 fits.
- Subtitle reading speed: check 20 characters per second before burning
Netflix caps English subtitles at 20 characters per second for adults, 17 for children. Check each cue's rate in a short script before a Sume caption render.
- NEXT_PUBLIC_ plus a Sume API key: why it ships to the browser
A NEXT_PUBLIC_ prefix inlines the value into client JavaScript at build time. Keep the Sume API key server-side, proxy via a route handler, rotate if it leaked.
- Poll a Sume job with AbortSignal.any in Node 26.10
Node 26.10.0 fixes AbortSignal.any() propagation. Here is a Sume job poll loop with a hard deadline and a caller cancel, using next_poll_after_seconds.
Written by Sume