httpx timeout: the 5-second default and slow AI API calls
httpx times out after 5 seconds of network inactivity by default. How to set connect, read, write and pool timeouts for slow AI API calls.

httpx's default timeout is 5 seconds: it raises a TimeoutException after 5 seconds of network inactivity. Change it with timeout= on a request or a Client, or pass httpx.Timeout(...) to set the connect, read, write and pool timeouts separately; timeout=None turns them off.
The httpx behavior comes from its Timeouts and Exceptions docs, read on 2026-09-28. The slow-call example is Sume's image generation endpoint, which can hold a request open for up to 30 seconds.
What are httpx's four timeouts?
Each one covers a different phase of the request and raises its own subclass of httpx.TimeoutException, so you can catch them together or apart.
| Timeout | What it limits | Exception |
|---|---|---|
| connect | Time to establish a socket connection to the host | httpx.ConnectTimeout |
| read | Wait for a chunk of data to be received, such as a chunk of the response body | httpx.ReadTimeout |
| write | Wait for a chunk of data to be sent, such as a chunk of the request body | httpx.WriteTimeout |
| pool | Wait to acquire a connection from the connection pool | httpx.PoolTimeout |
How do I set a timeout in httpx?
Pass a number to set all four at once, or an httpx.Timeout with a default and per-phase overrides. A timeout on the client becomes the default for every request it makes, and a per-request timeout= overrides it. AsyncClient works the same way, with await on each request.
import httpx
httpx.get("https://example.com/", timeout=10.0) # all four phases: 10 s
client = httpx.Client(timeout=httpx.Timeout(10.0, connect=60.0))
client.get("https://example.com/", timeout=None) # no timeout for this call
async def fetch():
async with httpx.AsyncClient(timeout=10.0) as aclient:
return await aclient.get("https://example.com/")Why do AI API calls raise httpx.ReadTimeout?
Because the server is still working and sends nothing back. A generation API that waits for the result before answering leaves the connection quiet, and the read timeout counts that silence.
Sume's POST /v1/images is an example: its mode defaults to sync on that route, and wait_timeout_seconds is 0–30 with a default of 30 (Image generation). With httpx's 5-second default, a request that Sume would answer at second 12 raises ReadTimeout at second 5. Two fixes:
- Set the read timeout above the server's wait: for a 30-second
wait_timeout_seconds, something likehttpx.Timeout(10.0, read=40.0). - Or send
mode: "async". Sume answers202with the job envelope right away, and you pollGET /v1/jobs/{id}/statuswith short requests that fit the default.
import os
import httpx
headers = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "poster-2026-09-28-001"}
body = {"model": "bytedance-seed/seedream-4.5",
"prompt": "a red panda astronaut, studio lighting",
"wait_timeout_seconds": 30}
with httpx.Client(timeout=httpx.Timeout(10.0, read=40.0)) as client:
r = client.post("https://api.sume.com/v1/images", headers=headers, json=body)
if r.status_code == 202: # still running: poll, don't resubmit
status_url = r.json()["data"]["status_url"]What happens to the job when httpx times out?
Your timeout only closes your own connection. Sume's docs say to keep the job id and recover with the jobs API instead of submitting duplicate paid work (Core workflow). If the timeout fired before you saw a job id, retry the submit with the same Idempotency-Key and the same body, so the retry returns the original job instead of billing a second one (Jobs and results).
Check the status code, not the body shape: 200 is the finished image response and 202 is a job to poll. Video generation API timeouts lists Sume's own wait caps and what happens when each one runs out.
Sources
Related posts
More in Developers
- Is POST idempotent? No, but PUT and DELETE are
No. HTTP defines GET, HEAD, OPTIONS, TRACE, PUT and DELETE as idempotent, but not POST or PATCH. An idempotency key makes a POST retry safe.
- Local vs remote MCP server: stdio or Streamable HTTP?
A local MCP server runs on your computer and talks over stdio; a remote one runs elsewhere and is reached by URL over Streamable HTTP. How to choose.
- Long polling vs short polling: what's the difference?
Short polling asks on a timer and gets an instant answer; long polling holds each request open until there's news or a timeout, then asks again.
- MCP error codes: what -32601, -32602, and -32001 mean
MCP error codes are JSON-RPC codes. What -32700 to -32603 mean, why -32001 is a client-side timeout, and how a failed tool call differs from both.
Written by Sume