Is GPT Image 2.5 really 50% faster? Time your own calls on Sume
OpenAI calls Flare 50% faster than Sunburst. That compares two models; it is not a promise for your prompt. A small Python timer for your own calls.

No single number answers this. OpenAI's announcement dated 2026-09-08 describes Flare as 50% faster than Sunburst, prioritising speed for everyday prompts, while Sunburst takes longer for precision. That compares two models. Your prompt, quality tier and size decide the real wait, so time your own calls: a 20-line script gives you a median in a few minutes.
What the sources say, and don't
The announcement states the Flare versus Sunburst comparison and says nothing you can budget as seconds. OpenAI's image generation guide adds that complex prompts may take up to 2 minutes. Neither source gives a median latency for a given size and quality, so any table of seconds you find elsewhere is a measurement of somebody else's prompts.
What you can say safely is directional: Flare for speed, Sunburst for precision, low quality before max, smaller before larger. That is also why the model choice and the quality tier belong in config.
| Claim | Source | Usable for budgeting? |
|---|---|---|
| Flare is 50% faster than Sunburst | OpenAI announcement | As a direction only |
| Complex prompts may take up to 2 minutes | OpenAI image generation guide | As a timeout upper bound |
| Sume waits 30 s, then returns a 202 job | Sume Image API docs | Yes, for your client design |
A small timer
The script posts the same prompt to both models with mode: "async" omitted, so the sync path is measured, and records wall-clock time. A 202 counts as slower than 30 seconds and is noted. Run it ten times per setting and look at the median, not the best run.
import os
import statistics
import time
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PROMPT = "A studio photo of a ceramic mug on a wooden table"
def timed(model: str, quality: str) -> float:
t = time.perf_counter()
r = requests.post(
"https://api.sume.com/v1/images",
headers=H,
json={"model": model, "prompt": PROMPT, "quality": quality},
timeout=60,
)
r.raise_for_status()
return time.perf_counter() - t if r.status_code == 200 else float("inf")
for model in ("openai/gpt-image-2.5", "openai/gpt-image-2.5-sunburst"):
runs = [timed(model, "medium") for _ in range(10)]
print(model, "median", round(statistics.median(runs), 1), "s", "slow:", runs.count(float("inf")))Read the result carefully
Network distance, your own upload of references, and queueing on shared capacity all show up in a wall-clock timer. Compare models inside the same run and at the same hour. If Flare's median is clearly lower than Sunburst's, that matches the direction of the announcement; if it is not, your prompt may not be the everyday case OpenAI describes. Note that '50% faster' can be read as a time ratio of 0.5 or 0.67, so do not hold either number against your results.
Sync waits cap at 30 seconds on Sume's Image API. Treat any 202 as the right place to switch to the job pattern in Jobs and results.
Sources
Related posts
More in Developers
- Is Kling O3 or LTX-2.5 on Sume? Read the catalog before you plan
New video models appear weekly in vendor feeds. Sume's catalog is one authenticated GET: this checks for Kling O3, LTX-2.5 and Wan 3.0 and reads capabilities.
- Why is my TTS job slow? Read the Sume job events timeline
GET /v1/jobs/{id}/events returns a sanitized lifecycle timeline. Use it to tell queue wait from generation time and failed webhook deliveries.
- Jupyter contact sheet: one prompt across four Sume image models
A notebook cell that sends one prompt to four Sume image models, tiles the results into a labeled contact sheet with Pillow, and prints each billed cost.
- Keep your Sora-style create_video() call: map it onto Sume
Sora's seconds, size and input_reference become duration, resolution plus aspect_ratio, and a first frame. Here is that map as a Python wrapper over Sume.
Written by Sume