Asynchronous request-reply pattern: how it works
In the asynchronous request-reply pattern, the server accepts work with a 202 and a status URL, and the client polls or takes a callback until it's done.

The asynchronous request-reply pattern lets a server take on long-running work without holding the HTTP connection open. The server answers the request at once with 202 Accepted, a job id and a status URL. The client polls that URL, or waits for a callback, until the job reaches a final state, then fetches the result from a separate URL.
The HTTP side comes from RFC 9110, the HTTP semantics standard. The worked example is Sume's generation job API, from its Jobs and results docs. Both were read on 2026-09-28.
Why does the pattern exist?
Because some work outlasts a request. RFC 9110 defines 202 Accepted for exactly this case: the request has been accepted for processing, but processing has not been completed. The code exists so a server can accept work for some other process without requiring the client's connection to persist until that process is done.
The RFC also notes that HTTP has no facility for re-sending a status code from an asynchronous operation. There is no second response, so the first one has to tell the client where to look: the RFC says it ought to describe the request's current status and point to a status monitor. That pointer is the status URL.
How does the asynchronous request-reply pattern work?
Five parts, in the order a client meets them:
- Accept: the server validates the request, stores a durable job and replies at once with the job id and the URLs to follow.
- Poll: the client reads the status URL on an interval, ideally one the server suggests, until the job is terminal.
- Fetch: the client reads the result from its own URL, which only answers once the job has completed.
- Callback (optional): the server calls the client's URL when the job ends. Polling stays available as the backup.
- Safe retry: a retried submit carries an idempotency key, so a lost reply doesn't start the work twice.
What does the pattern look like in a real API?
Each part maps to a field or route. In Sume's job API, a 2xx on submit means the job exists and paid work is in flight; it does not mean the job finished.
| Pattern part | Sume's job API |
|---|---|
| Accept | Default async mode returns 202 with the job envelope and polling URLs; every mode returns the job id first |
| Status URL | status_url (GET /v1/jobs/{id}/status); poll until terminal is true |
| Poll interval | Honor next_poll_after_seconds when present, back off otherwise |
| Result URL | result_url, only for completed jobs; any other state, failed or canceled included, answers 409 job_not_completed, so read a failure off the job record |
| Final states | completed, failed, canceled |
| Callback | mode: "webhook" delivers only job.completed, job.failed and job.canceled |
| Safe retry | The same Idempotency-Key returns the original job instead of billing a second one |
Why not keep the request open until the work is done?
Because something between client and server will close it first. Sume caps the blocking wait on a submit at 30 seconds, and its docs note that every edge eventually closes a request held open with nothing transferring. Video jobs routinely outlast that. Sync vs async video generation covers the wait modes, and 504 gateway timeout from an API covers what a closed request looks like.
What should a client do when the reply is lost?
Ask the status URL before doing anything else, because closing the connection doesn't stop the work. On Sume, a client-side timeout does not cancel the job: it keeps running and still bills, so keep the job id and recover with the jobs API instead of submitting duplicate paid work. If the submit itself got no answer, retry it with the same Idempotency-Key and the same body. 504 Gateway Timeout from an API walks through each case, and Idempotency keys for AI video APIs covers the key rules.
Sources
Related posts
More in Developers
- asyncio Semaphore: limit concurrent API jobs in Python
An asyncio Semaphore caps how many coroutines run a block at once. For paid API jobs, hold it from submit to the final status and size it to your limit.
- Bearer token vs API key: what's the difference?
An API key is a kind of credential; Bearer is a way to send one, in the Authorization header. A key can travel as a bearer token, as OAuth tokens do.
- Circuit breaker pattern in Python for AI API calls
A circuit breaker stops calling a failing API: after repeated failures it opens and fails fast, then lets a trial call through. A Python version.
- How to create an SRT file from text: time it with TTS
An SRT file needs a start and end time for every line. Voice your text with TTS word timings, then write each timed sentence as a numbered block.
Written by Sume