Python urllib: POST /v1/images, 200 or 202, after Imagen 4 Fast
imagen-4.0-fast-generate-001 ended Aug 17. A stdlib Python call to Sume's image route that reads the status code, then polls the job when the answer is 202.

Check the HTTP status, not the body: 200 means data holds image URLs, and 202 means the body is a job envelope you must poll. Google lists imagen-4.0-fast-generate-001 among the Imagen 4 ids shut down on August 17, 2026, so a Python caller moving to Sume's POST /v1/images has to handle both answers.
Why both answers exist
/v1/images waits up to 30 seconds by default. Most models finish inside that and return 200. A slow request, such as 4K, high quality or a large n, can run out of wait and degrade to 202, and mode: "async" or "webhook" always returns 202.
| Status | Body | Next step |
|---|---|---|
| 200 | created, data[] with url and media_type, usage | Use the URLs |
| 202 | data.job, status_url, result_url | Poll status, then read the result |
| 402 | error insufficient_credits | Nothing started |
| 429 | rate_limited or queue_full | Back off, same key |
The call
The sample uses only urllib. Non-2xx answers raise HTTPError, which is what you want for a 402 or 400. On 202 it polls /v1/jobs/{id}/status until terminal is true, honors the next_poll_after_seconds hint with a 2-second floor, then reads /result and pulls each artifact's url.
The model id is google/nano-banana-2.1, the id in Sume's image docs. If you still want an Imagen variant, read its id from GET /v1/images/models; the sample does not assume one.
import json
import os
import time
import urllib.request
BASE = os.environ.get("SUME_BASE", "https://api.sume.com")
HEADERS = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}
def call(method, path, body=None, extra=None):
req = urllib.request.Request(BASE + path, method=method, headers={**HEADERS, **(extra or {})},
data=json.dumps(body).encode() if body else None)
with urllib.request.urlopen(req, timeout=40) as res:
return res.status, json.load(res)
status, body = call("POST", "/v1/images", {"model": "google/nano-banana-2.1", "prompt": "red kettle on a white shelf"},
{"Idempotency-Key": "kettle-0001"})
if status == 200:
urls = [img["url"] for img in body["data"]]
else: # 202: the 30 s wait expired, or you asked for async; the body is the job envelope
job_id = body["data"]["job"]["id"]
while True:
_, st = call("GET", f"/v1/jobs/{job_id}/status")
if st["data"]["terminal"]:
break
time.sleep(max(2, st["data"].get("next_poll_after_seconds") or 2))
_, result = call("GET", f"/v1/jobs/{job_id}/result")
urls = [a["url"] for a in result["data"]["result"]["artifacts"]]
print(urls)Gaps
There is no deadline on the poll, no handling of a failed job (read the job's error), and no 429 backoff. The Idempotency-Key stays the same across a retry, so a repeated submit adopts the first job. Price differs from the retired model, and usage.cost on a 200 is the amount billed.
Sources
Related posts
More in Developers
- queue_full 429 on a Sume submit: the reservation is released
A 429 queue_full releases or refunds the failed admission's reservation. Check refunded_usd_micros in /v1/usage, then retry with the same Idempotency-Key.
- Can I submit 100 AI video jobs at once? Queue limits by plan
Accepted capacity is slots plus queue: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Submit 100 at once and 94, 76, 52 or 0 get 429 queue_full.
- Read the motion clip length with video inspect before Kling duration
Kling motion control on Sume reserves money from the duration_seconds you declare. Probe the reference clip with video inspect first, so the number is measured.
- Reconcile Sume jobs after a deploy or outage: poll what is open
After downtime, read status for every job your own table still shows as open, honor terminal and result_ready, and never resubmit. Python with sqlite.
Written by Sume