AWS Lambda's 3 s default vs POST /v1/images' 30 s sync wait
POST /v1/images waits up to 30 seconds by default, but a Lambda defaults to 3. Send mode async with an Idempotency-Key and return the job id instead.

Set "mode": "async" on POST /v1/images when you call it from AWS Lambda. Lambda's timeout page gives a default of 3 seconds and a maximum of 900 seconds, while the Sume image route, unless you say otherwise, waits for the result in sync mode for up to 30 seconds. A function left on its default timeout can be stopped in the middle of that wait, with the image paid for and nobody holding the answer.
The fix is small: ask for the job, not the picture. In async mode the route returns a 202 with a job envelope that includes a status_url. Your function stores the job id, returns, and a later invocation, a webhook or a poller collects the result.
How the image route decides what to return
According to the image model docs, the route is synchronous by default to match the OpenRouter image API shape. It returns 200 with the images if the job finishes within the wait, 502 with an error envelope if the job fails during the wait, and 202 with the job envelope if the wait runs out or if you chose async or webhook. Slow settings such as 4K, high quality or a large n can fall back to the 202 on their own, so even a long Lambda timeout needs code that branches on the status code.
| Setting | Value | Source |
|---|---|---|
| Lambda default timeout | 3 seconds | AWS docs |
| Lambda maximum timeout | 900 seconds | AWS docs |
| POST /v1/images default mode | sync, wait_timeout_seconds 30 | Sume image docs |
| Result in time | 200 with images | Sume image docs |
| Failed in the wait | 502 with error envelope | Sume image docs |
| Timed out, or async / webhook | 202 with job envelope | Sume image docs |
Steps
- Add
"mode": "async"to the request body. Keepmodelandpromptas in the sync call. - Send an
Idempotency-Keyderived from your order, so a Lambda retry returns the original job (idempotency_hit: true) instead of creating a second paid one. - Use a client timeout shorter than the function timeout, so you return a clean error instead of being killed.
- Persist
data.job.idanddata.status_url. Collect the result withGET /v1/jobs/{id}/resultonce status isCOMPLETED, or receive it by webhook.
Handler
This handler uses only the standard library. It reads the key from the environment and the order key from the event. It was run against a local stand-in for the API, not the live service.
import json
import os
import urllib.request
BASE = os.environ.get("SUME_BASE", "https://api.sume.com")
def handler(event, context=None):
body = {"model": "sume/auto", "prompt": event["prompt"], "mode": "async"}
req = urllib.request.Request(
f"{BASE}/v1/images",
data=json.dumps(body).encode(),
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": event["order_key"],
},
)
with urllib.request.urlopen(req, timeout=2) as res:
doc = json.load(res)
job = doc.get("data", {}).get("job", {}).get("id")
return {"status": res.status, "job_id": job, "poll": doc["data"]["status_url"]}
if __name__ == "__main__":
print(handler({"prompt": "paper boat on a pond", "order_key": "order-1042-v1"}))What Sume does not do
Sume does not shorten the sync wait to fit your function, and it does not tell the route what your platform timeout is. If you send the default body from a 3 second function and the platform stops you, the job may already exist; retrying with the same Idempotency-Key returns it rather than creating another. Read the jobs and results docs for the poll and result routes.
Sources
Related posts
More in Developers
- Lambda get_remaining_time_in_millis: stop polling a Sume job in time
A Lambda that polls a Sume job should check get_remaining_time_in_millis before each sleep and return the job id so a later invocation can resume. Python.
- Largest square GPT Image 2.5 takes on Sume: 2880x2880
2880x2880 is exactly 8,294,400 pixels, the top of the GPT Image 2.5 custom-size box on Sume. 2896x2896 is over. 4K UHD lands on the same cap.
- List your Sume Formats in Python and keep the video ones
Page through GET /v1/formats with next_cursor, keep io.output_kind video, and print each vanity_invoke_url. A runnable snippet with field caveats.
- List Sume jobs for one run_id and count them by status (curl and jq)
GET /v1/jobs filters by run_id, scope and status and returns up to 100 rows. A short curl and jq script gives a count per status for one run.
Written by Sume