GPT Image 2.5 can take two minutes: handle the 202 job on Sume

OpenAI says complex GPT Image 2.5 prompts can take up to two minutes. Sume blocks 30 seconds, then returns a 202 job. Submit async, poll, and fetch the result.

5 min readSume
All posts

OpenAI's image generation guide warns that complex prompts may take up to 2 minutes. Sume's POST /v1/images waits at most 30 seconds, so a slow GPT Image 2.5 request returns 200 with the image if it finishes inside the window and 202 with a job envelope if it does not. Check the status code, not the body shape, and for anything that could be slow submit with mode: "async" and poll the job.

Which requests run long

The Image API docs name the slow configurations: 4K, high quality, and a large n. On GPT Image 2.5 the quality levels are low through max, so treat high, xhigh and max as the slow end, and OpenAI also says complex prompts can run long. A single low image at 1024x1024 will usually come back inside the window; do not build a pipeline on usually.

Status codes from POST /v1/images (read 2026-10-04)
StatusMeaningWhat to do
200Finished within the 30 s waitRead data[].url
202Job envelope: still running, or you asked for async or webhookPoll status_url, then fetch result_url
400 unsupported_parameterA parameter the model does not listRead the model's descriptors
502Generation failed (not billed)Retry the request

Submit async from the start

Sending mode: "async" skips the blocking wait entirely and always returns the 202 envelope, which is the pattern Jobs and results recommends for anything that can outlast 30 seconds. The wait lives in your client, so its timeout can be minutes without holding an HTTP request open. Reuse the same Idempotency-Key if you retry the submit, so you do not bill a second job.

Do not submit a new paid job for the same intent just because the wait timed out; the docs say wait exhaustion is not an admission failure.

import os
import time
import requests

API = "https://api.sume.com"
H = {
    "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
    "Idempotency-Key": "poster-hero-001",
}
body = {
    "model": "openai/gpt-image-2.5-sunburst",
    "prompt": "A dense festival poster with a 3x3 grid of illustrated food stalls",
    "quality": "xhigh",
    "mode": "async",
}
r = requests.post(f"{API}/v1/images", headers=H, json=body, timeout=60)
job = r.json()["data"]
while True:
    s = requests.get(job["status_url"], headers=H, timeout=30).json()
    s = s.get("data", s)
    if s.get("terminal"):
        break
    time.sleep(5)
print(requests.get(job["result_url"], headers=H, timeout=30).json())

Or take a webhook

For batches, mode: "webhook" with a webhook_url delivers the terminal event to your server so no process sits polling. Verify the signature scheme described in Webhooks and reject requests when your signing secret is empty.

What to tell your users

Show a progress state after about 10 seconds rather than an error. A generation that fails is not billed. Client disconnects count as failed generations, so closing the tab on a sync call does not charge you, but it also throws the result away. Async keeps it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume