Recraft V4.1 Flash: median 1.3 s, p95 1.8 s. Set timeouts from p95
Recraft quotes a median of about 1.3 seconds and a p95 of 1.8 seconds for V4.1 Flash. How to turn latency claims into timeouts and polling for image APIs.

Recraft reports a median latency of about 1.3 seconds and a p95 of 1.8 seconds for V4.1 Flash. Set your client timeout from the p95 plus transport margin, not from the median, because the median says nothing about the slow tail; by definition about one request in twenty runs longer than the p95.
What Recraft published
The Recraft announcement dated September 23 gives the median and p95 figures, and describes a Refine button that upscales to 2048 by 2048 with V4.1 Pro at Subtle or Moderate strength. The model is available in Recraft Studio and through the API.
| Item | Value |
|---|---|
| Median latency | About 1.3 s |
| p95 latency | 1.8 s |
| Refine output | 2048x2048 with V4.1 Pro |
| Refine strengths | Subtle, Moderate |
| Availability | Recraft Studio and API |
From a claim to a timeout
A vendor latency number describes the model run under the vendor's conditions. Your request also pays for network time, queueing and any media handling around the call. A defensible starting rule is to set the timeout to several multiples of p95, then observe your own tail and tighten it.
Sequential totals are medians only. One hundred draft images in a row would take about 130 seconds at the median, which is a planning figure, not a promise. Parallel requests change the picture again once rate limits or concurrency caps apply.
What changes when the call goes through Sume
A vendor's model latency is one part of a Sume job. Sume accepts valid generation requests as queued, and workspace concurrency limits apply when workers move jobs into processing. For POST /v1/images, the request blocks for up to 30 seconds and returns 200 with the image if it finishes in time, or 202 with a job envelope if it does not. Check the status code, not the body shape.
A 30-second wait budget is therefore comfortable for a model that finishes in under two seconds, but your code should still handle 202. Do not resubmit a paid request because a local timer expired; poll the job instead, or reuse the same Idempotency-Key when retrying the submit.
curl -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: draft-0001" \
-d '{"model": "sume/auto", "prompt": "flat vector icon of a paper plane", "mode": "async"}'A draft-then-refine loop
The Flash plus Refine pairing suggests a workflow: generate many cheap, fast drafts, pick one, then spend on the finishing pass. The same shape works with any pair of models. Log your own median and p95 over a few hundred calls so that the next vendor claim you read has something to be compared against.
Sources
Related posts
More in Developers
- Reproduce the same AI voiceover later: model id, voice, settings
To redo a narration line months later you need the model id, voice, language, format and settings. Sume's completed TTS job records them. A short routine.
- Shopify rejects file names ending in thumb, icon or large
Shopify file uploads reject names ending in pico, icon, thumb, testing, small, compact, medium, large or grande. A Python rename step for batch outputs.
- Shopify image limits: 20 MB, 25 megapixels vs Sume image outputs
Shopify accepts product images up to 20 MB and 25 megapixels in JPEG, PNG, WEBP, HEIC or GIF. How that lines up with Sume image model sizes and formats.
- Sign MCP requestState for paid render approvals: user, TTL, digest
MCP says requestState is attacker-controlled. For a paid render approval, bind it to the user, an expiry and an argument digest with an HMAC. Python included.
Written by Sume