GPT Image 2.5 slow? OpenAI says jpeg is faster than png
OpenAI says jpeg is faster than png and complex prompts can take 2 minutes. How to set output_format on Sume and handle the 202 that follows.

If GPT Image 2.5 calls feel slow, set output_format to jpeg. OpenAI's image generation guide says that jpeg is faster than png and that you should prioritize it if latency is a concern; it also warns that complex prompts may take up to 2 minutes. On Sume the same field takes png, jpeg or webp, and a call that outlasts the 30-second wait returns 202 with a job envelope.
What the vendor says
Two statements in OpenAI's guide matter here, both read on 2026-10-08. First, png is the default format, and jpeg and webp are options; using jpeg is faster than png. Second, under limitations, complex prompts may take up to 2 minutes. The guide gives no per-format timings, so this post does not either.
| Lever | Effect | Source |
|---|---|---|
| output_format: jpeg | Faster than png, per OpenAI | OpenAI image guide |
| quality: low | Recommended for quick drafts | OpenAI image guide |
| Complex prompt | Up to 2 minutes | OpenAI image guide |
| Slow call on Sume | 200 within 30 s, else 202 job | docs.sume.com/models/images |
On Sume
Send the field on POST /v1/images. Image price on Sume is per image, so the format does not change the bill; check usage.cost if you want to confirm for your own calls. Note that jpeg has no alpha channel, so it cannot carry background: "transparent", which needs png or webp.
Sume does not serve output_compression in v1, so you cannot request a particular jpeg quality; a request with that field returns 400 unsupported_parameter.
{
"model": "openai/gpt-image-2.5",
"prompt": "wide hero photo of a mountain cabin at dusk",
"quality": "medium",
"output_format": "jpeg",
"aspect_ratio": "16:9"
}Plan for the slow case anyway
Faster formats and lower quality reduce the chance of a 202; they do not remove it. Treat the status code as the contract: 200 carries the images, 202 carries status_url and result_url. The jobs docs describe polling with exponential backoff and stopping when the job is terminal.
For batches, use mode: "async" from the start and poll in a worker. That removes the 30-second window from the design.
- Use jpeg for photos you will publish as jpeg anyway.
- Use png or webp when you need transparency or lossless edges.
- Use low quality for layout checks; switch to high for the final.
- Log the elapsed time per call so you can see your own numbers.
Measure it yourself
Run ten identical prompts, five as png and five as jpeg, at the same quality and ratio, and log the wall-clock time per call. Sume's docs promise a 30-second bounded wait, not a latency figure, so your own numbers are the only ones that matter for your region and your prompt mix. If the jpeg runs are not faster for you, the format is not your bottleneck; look at quality tier and prompt complexity next.
What to do with a slow job
If a call returns 202, do not resubmit it. The first job is still running and will bill when it completes; a second submit creates a second job. Poll the status URL from the envelope, honor next_poll_after_seconds when it is present, and use exponential backoff otherwise. A client timeout stops your wait, not the job.
Sources
Related posts
More in Developers
- Jupyter: move Sora cells to Sume, where a cell rerun is a retry
Re-running a notebook cell resubmits the request. Hold one Idempotency-Key per take in a variable so Sume returns the first video job, and show the file inline.
- Kotlin: submit and poll a Sume video job with java.net.http
A Kotlin port of a Sora videos call: POST /v1/videos with an Idempotency-Key, poll every 30 seconds until done, with the JDK client and kotlinx.serialization.
- Kubernetes CronJob that submits a nightly 30-second Wan 3.0 clip
A CronJob manifest using curlimages/curl and a Secret: one dated Idempotency-Key per night, concurrency forbidden, and a month of reserves at each resolution.
- AWS Lambda's 3 s default vs POST /v1/images' 30 s sync wait
POST /v1/images waits up to 30 seconds by default, but a Lambda defaults to 3. Send mode async with an Idempotency-Key and return the job id instead.
Written by Sume