gpt-image-2.5 partial image streaming vs Sume's 400
OpenAI streams 0 to 3 partial images for gpt-image-2.5. Sume returns 400 streaming_not_supported for stream: true. What to send instead: a job and a poll.

OpenAI's image generation guide says gpt-image-2.5 can stream partial image previews, between 0 and 3 partial images per request. Sume does not serve that: sending stream: true to POST /v1/images returns 400 streaming_not_supported, because every catalog row reports supports_streaming: false in v1.
What exactly does OpenAI offer?
The OpenAI guide read on 2026-10-02 lists streaming with partial image previews for both gpt-image-2.5-sunburst and gpt-image-2.5-flare, configured as 0 to 3 partial images. It also notes that complex prompts can take up to two minutes, which is the reason previews exist: they give a user something to look at while the final render finishes.
That is a client-experience feature. It does not change the final image, its quality, or its price, so the question for an integration is whether your UI needs progressive previews or simply a reliable result.
What does Sume do instead?
The Image API states that Sume does not serve native SSE in v1. The stream field is in the schema so that clients can adopt it without a code change later, but at runtime it is rejected. mode: "subscribe" is not a progress stream either; it is an alias of sync and buys one bounded 30-second wait.
What you get is a job model. POST /v1/images blocks for up to 30 seconds. If the generation is still running, or you send mode: "async" or mode: "webhook" with a webhook_url, Sume returns 202 with a job envelope, and you read the images from the standard job result endpoint.
How do the two approaches line up?
The table maps each need to the Sume behaviour documented today.
| You want | OpenAI gpt-image-2.5 | Sume today |
|---|---|---|
| Progressive previews | 0 to 3 partial images via streaming | Not available; stream: true is a 400 |
| Result without holding a connection | Long request or background flow | mode: async or webhook returns 202 and a job |
| Terminal notification | Your own handling | Signed job.completed, job.failed, job.canceled webhook |
| Final image delivery | Image data in the response | Sume-hosted signed URL in data[].url |
How do I handle a 200 versus a 202?
Check the status code, not the body shape. A 200 is the image response with data[].url; a 202 is the job envelope with status_url and result_url. Slow configurations, such as 4K, high quality or a large n, are the ones most likely to degrade to 202, per the docs.
A minimal client therefore branches on the status code, polls GET /v1/jobs/{id}/status when it gets 202, and then fetches GET /v1/jobs/{id}/result. A webhook is the better path for a server, with polling kept as the fallback, as the webhooks page recommends.
When is this a dealbreaker?
If your product shows a blurry-to-sharp preview inside a chat window, you will not get that from Sume's image route today, and you should say so in your design rather than fake it. If you generate images in a pipeline, in a batch, or behind a queue, previews add nothing, and the job model is a better fit.
Sume is honest about the gap in its docs, and it promises only that the field exists. Do not build against a streaming date; there is none in the documentation I read. See long requests and 202 jobs for the handling pattern.
Sources
Related posts
More in Comparisons
- gpt-image-2.5 rate tiers (5 to 250 IPM) vs Sume plan queue
OpenAI limits gpt-image-2.5 Flare from 5 images per minute at Tier 1 to 250 at Tier 5. Sume limits processing concurrency by plan and queues the rest.
- H3 Max Recast vs Genjutsu: which person swap to call on Sume
Sume lists two person-swap video rows. Recast: 1-4 people in a 5-30 s clip at 768p or 1080p. Genjutsu: 1-8 images at 480p or 720p. How to choose.
- Hedra 402 INSUFFICIENT_BALANCE vs Sume insufficient_credits
Hedra returns 402 INSUFFICIENT_BALANCE until you add funds; Sume returns 402 insufficient_credits. How each wallet check behaves and how to preflight a job.
- Hedra Avatar needs start frame and audio; Sume takes a handle and scri
Hedra Avatar generates from a start frame plus an audio track, up to 10 minutes. Sume's talking-video takes an avatar handle and a script, up to 60 seconds.
Written by Sume