Recraft V4.1 Flash: median 1.3 s, p95 1.8 s. Set timeouts from p95

Recraft quotes a median of about 1.3 seconds and a p95 of 1.8 seconds for V4.1 Flash. How to turn latency claims into timeouts and polling for image APIs.

4 min readSume
All posts

Recraft reports a median latency of about 1.3 seconds and a p95 of 1.8 seconds for V4.1 Flash. Set your client timeout from the p95 plus transport margin, not from the median, because the median says nothing about the slow tail; by definition about one request in twenty runs longer than the p95.

What Recraft published

The Recraft announcement dated September 23 gives the median and p95 figures, and describes a Refine button that upscales to 2048 by 2048 with V4.1 Pro at Subtle or Moderate strength. The model is available in Recraft Studio and through the API.

Recraft V4.1 Flash figures (read 2026-10-03)
ItemValue
Median latencyAbout 1.3 s
p95 latency1.8 s
Refine output2048x2048 with V4.1 Pro
Refine strengthsSubtle, Moderate
AvailabilityRecraft Studio and API

From a claim to a timeout

A vendor latency number describes the model run under the vendor's conditions. Your request also pays for network time, queueing and any media handling around the call. A defensible starting rule is to set the timeout to several multiples of p95, then observe your own tail and tighten it.

Sequential totals are medians only. One hundred draft images in a row would take about 130 seconds at the median, which is a planning figure, not a promise. Parallel requests change the picture again once rate limits or concurrency caps apply.

What changes when the call goes through Sume

A vendor's model latency is one part of a Sume job. Sume accepts valid generation requests as queued, and workspace concurrency limits apply when workers move jobs into processing. For POST /v1/images, the request blocks for up to 30 seconds and returns 200 with the image if it finishes in time, or 202 with a job envelope if it does not. Check the status code, not the body shape.

A 30-second wait budget is therefore comfortable for a model that finishes in under two seconds, but your code should still handle 202. Do not resubmit a paid request because a local timer expired; poll the job instead, or reuse the same Idempotency-Key when retrying the submit.

curl -X POST "https://api.sume.com/v1/images" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: draft-0001" \
  -d '{"model": "sume/auto", "prompt": "flat vector icon of a paper plane", "mode": "async"}'

A draft-then-refine loop

The Flash plus Refine pairing suggests a workflow: generate many cheap, fast drafts, pick one, then spend on the finishing pass. The same shape works with any pair of models. Log your own median and p95 over a few hundred calls so that the next vendor claim you read has something to be compared against.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume