POST /v1/images returns 200 or 202: one Python handler for Sume
Sume's image route blocks up to 30 seconds and returns 200 with images, or 202 with a job envelope. One handler that covers both and never double-submits.

POST /v1/images on Sume waits for up to 30 seconds. If the image finishes in that budget you get 200 and the images in data; if not, you get 202 with a job envelope (data.job.id, status_url, result_url), and the same envelope if you send mode: "async" or a webhook. Your client must handle both outcomes, and must treat 202 as success, not as a retry signal.
The shape of the 200 and the 202 are different, and the 202 body is the standard job shape, not the image body.
Two outcomes, two shapes
The 200 body mirrors the response format in the image docs; the 202 body is the job envelope. After 202, use the standard job endpoints.
| Status | Body | Next step |
|---|---|---|
| 200 | data[].url (Sume-hosted), usage.cost | Use the URLs |
| 202 | data.job.id, status_url, result_url | Poll GET /v1/jobs/{id}/status, then /result |
| 402 | insufficient_credits | Add funds or cheaper request |
| 429 | rate_limited or queue_full | Back off; same idempotency key |
Why 202 is not a failure
A 2xx means the job exists. If the wait budget ended, the job is still running and still billing. Do not submit again for the same intent. If you must retry the submit itself, use the same Idempotency-Key so the retry returns the original job.
Most catalog models finish inside the budget, so the 202 branch is the one that rarely runs and therefore the one most likely to be broken in production. Test it by forcing mode: "async" in a staging check.
The handler
This uses only the standard library. It reads the poll fields the docs name, terminal and result_ready, and unwraps a data envelope when one is present. Adjust the unwrap if your live response differs.
import json, os, time, urllib.request
BASE = "https://api.sume.com/v1"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}
def call(method, path, body=None, extra=None):
req = urllib.request.Request(BASE + path, method=method, headers={**HEAD, **(extra or {})},
data=json.dumps(body).encode() if body else None)
with urllib.request.urlopen(req, timeout=60) as resp:
return resp.status, json.load(resp)
def make_image(prompt, key):
status, body = call("POST", "/images", {"model": "bytedance-seed/seedream-4.5", "prompt": prompt},
{"Idempotency-Key": key})
if status == 200:
return [i["url"] for i in body["data"]]
job_id = body["data"]["job"]["id"]
while True:
_, st = call("GET", f"/jobs/{job_id}/status")
st = st.get("data", st)
if st.get("terminal"):
break
time.sleep(float(st.get("next_poll_after_seconds") or 3))
_, res = call("GET", f"/jobs/{job_id}/result")
res = res.get("data", res)
return [a["url"] for a in res["result"]["artifacts"]]
print(make_image("a red panda astronaut, studio lighting", "panda-001"))
Edge cases
A failed job returns 409 job_not_completed from /result, so read the failure from GET /v1/jobs/{id} instead. Disconnecting your client does not cancel the job. And n and the other parameters vary by model: read GET /v1/images/models/{id}/endpoints rather than guessing, since Sume returns 400 unsupported_parameter for a field the model does not list.
Testing the 202 branch
Add one integration test that submits with mode: "async" and asserts that your code returns the same artifact list as the synchronous path. Add another that kills the process after the 202 and resumes from the stored job id. These two tests catch most production bugs in image integrations, because the happy path rarely exercises the job branch.
Finally, log which branch each request took. If the 202 share rises over a week, that is a signal that the model got slower or your prompts got heavier, and it is cheaper to notice from a counter than from user complaints. Store the job id on both branches so support can look up any request, including the ones that returned inline, without asking the user for a screenshot or a timestamp.
Sources
Related posts
More in Developers
- Probe a finished video before upload: duration, size and aspect
Run video inspect with frames false to read a render's duration, size and frame rate before posting. Check it against the 3-minute Shorts limit.
- URL or QR code in an AI avatar video: say it, caption it, or link it
An avatar clip cannot reliably carry a QR code or a long URL. Three Sume-supported options: spoken words, authored caption cues, or the text beside the video.
- Python asyncio.Semaphore sized to a Sume plan's job capacity
Free holds 6 paid jobs at once, Pro 24, Startup 48, Scale 120. A Semaphore of that size keeps a 60-job batch from hitting 429 queue_full.
- Python: check Omni Flash 1.1 limits against /v1/videos/models first
Google's Omni runs a sync call and extends in 10-second steps up to 40 seconds. Sume's gemini-omni-flash-1.1 takes 3 to 10 seconds per job. A preflight check.
Written by Sume