GPT Image 2.5 4K returns 202: a Python poll loop that works

Slow ChatGPT Image 2.5 calls on Sume return 202 and a job envelope after 30 seconds. Python that handles 200 and 202 without paying twice.

5 min readSume
All posts

POST /v1/images waits up to 30 seconds and returns 200 with the images, or 202 with a job envelope if the generation is not done; read the status code, not the body shape. A 4K max ChatGPT Image 2.5 call is a likely 202, and the fix is to poll the job instead of sending the paid request again.

The envelope carries status_url and result_url. Poll the first until terminal is true, then fetch the second when result_ready is true.

The loop

This script sends the request, handles both outcomes, and honors next_poll_after_seconds when the status response includes it. Set SUME_API_KEY in your environment.

import os, time, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {"model": "openai/gpt-image-2.5", "prompt": "Alpine lake at dawn",
        "image_size": "3840x2160", "quality": "high"}

def main():
    r = requests.post("https://api.sume.com/v1/images", headers=H, json=body, timeout=60)
    if r.status_code == 200:
        return print([i["url"] for i in r.json()["data"]])
    env = r.json()["data"]
    wait = 2
    while True:
        s = requests.get(env["status_url"], headers=H, timeout=30).json()
        s = s.get("data", s)
        if s.get("terminal"):
            break
        time.sleep(s.get("next_poll_after_seconds", wait))
        wait = min(wait * 2, 15)
    if not s.get("result_ready"):
        raise SystemExit(f"job ended without a result: {s.get('sume_status')}")
    print(requests.get(env["result_url"], headers=H, timeout=30).json())

main()

What each outcome costs

A 200 response reports the cost in usage.cost. A 202 job is billed when it completes, and a failed or cancelled job is not charged. Per the jobs docs, do not submit the original paid request again only because a local process timed out.

Outcomes of POST /v1/images (read 2026-10-07)
StatusMeaningWhat to do
200Images in the response, Sume-hosted URLsRead data[].url and usage.cost
202Job envelope, still runningPoll status_url until terminal
400 unsupported_parameterA field the model does not listFix the request; nothing is billed
502Generation failedNo charge; retry with the same inputs

Failure handling

A loop that waits forever is a bug. Put a ceiling on total waiting time, such as ten minutes, and stop with a clear message if it passes. A job that fails is terminal, and the loop above raises with the sume_status so you can see which end state it reached. Do not resubmit on a timeout of your own; check the job first, because it may be still running and will bill if it completes.

For many jobs, switch from polling to a webhook. Send mode: "webhook" with a webhook_url and your receiver gets a terminal event. If you verify a signature, refuse to run when the secret is empty.

  • Cap total wait time.
  • Log the job id before the first poll.
  • Use backoff and honor next_poll_after_seconds.

Reducing the odds of a 202

The docs name 4K, high quality and large n as the settings that most often exceed the wait. Cut n to one per call and run calls in parallel, or choose a lower tier for drafts. For a priced ladder at 3840x2160, see the 4K price post.

Send mode: "async" when you know the job is slow. It returns the envelope immediately and avoids holding a connection open for 30 seconds.

If you need to run several 4K jobs at once, start them all in async mode and collect the envelopes, then poll them in one loop. Log the job id and the quality tier together so a long-running job can be matched to its cost.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume