AI image API timeouts in Python: httpx above 30 s, 200 or 202
Sume's image endpoint holds a request up to 30 seconds, then returns 202 with a job. A runnable httpx example with a 40 second client timeout.

Set your HTTP client timeout above 30 seconds when you call Sume's image endpoint, and branch on the status code. POST /v1/images blocks for up to 30 seconds and answers 200 with the image; if the generation is still running when that budget ends, or if you ask for mode: "async", it answers 202 with a job envelope instead (Image API). A default client timeout of 5 to 10 seconds will cut the call before either answer arrives.
Runnable example
This uses httpx with a 40 second timeout. It prints the image URLs on 200, and the result URL on 202, so you can poll or take a webhook later. The key comes from the environment and the script refuses to run without it.
import asyncio, os
import httpx
async def main():
key = os.environ.get("SUME_API_KEY", "")
if not key:
raise SystemExit("set SUME_API_KEY")
body = {
"model": "black-forest-labs/flux.2-pro",
"prompt": "a ceramic mug on a pale oak table, soft morning light",
"aspect_ratio": "4:3",
}
async with httpx.AsyncClient(timeout=40.0) as client:
r = await client.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {key}"},
json=body,
)
if r.status_code == 200:
for item in r.json()["data"]:
print(item["url"])
elif r.status_code == 202:
print(r.json()["data"]["result_url"])
else:
raise SystemExit(f"{r.status_code}: {r.text}")
asyncio.run(main())Which requests slip to 202
Slow configurations are the ones most likely to degrade to 202: 4K output, high quality, and a large n. If you know a call is slow, send mode: "async" or mode: "webhook" with a webhook_url and stop holding a connection (jobs and results). wait_timeout_seconds accepts 0 to 30 on this route.
| Mode | Response | Client holds the line |
|---|---|---|
| sync (default) | 200 image, or 202 if it overruns | Up to 30 s |
| async | 202 job envelope | No |
| webhook | 202 plus a callback to your URL | No |
| subscribe | Same as sync: one bounded wait | Up to 30 s |
Retries
Do not retry blindly on a timeout: the generation may still be running and billing on completion. Reuse the same Idempotency-Key only for the same operation and payload when you retry.
Sources
Related posts
More in Developers
- AI video person swap rejects long takes: find shots over 15 s
H3 Max Recast on Sume needs 5 to 30 seconds with no single shot over 15. Find the long takes with ffmpeg scene detection, then cut them before you pay.
- AI voice reads Spanish with an English accent: two causes to check
Spanish read with an English accent usually means the language field was omitted or the voice is tagged for another language. How Sume's TTS 409 works.
- Are Claude Code mods safe with a Sume API key in your env?
Claude Code mods run unsandboxed and can read env vars and settings files. What that means for a Sume API key, the CLI config file and an OAuth session.
- Canary 10% of video jobs to Sume before cutover: sticky bucketing
Moving video traffic off a shut-down API: hash a stable key into a percent bucket so each customer stays on one backend, and raise Sume's share in steps.
Written by Sume