Sume SDK retries: which requests it repeats and how long it waits
createSumeClient retries 408, 429 and 5xx up to maxRetries (default 2), honours retry-after up to 60 seconds, and replays a POST only with an Idempotency-Key.

createSumeClient retries a request on 408, 429 and any 5xx, up to maxRetries times, and the default is 2. It honours the retry-after header and caps the wait at 60 seconds, and it adds about 20 percent jitter to every backoff. A POST is retried only when it carries an Idempotency-Key, because a repeat without one could start a second paid job.
Those defaults are conservative on purpose. Two retries cover a short blip, and the cap of 60 seconds stops a long retry-after from freezing a request handler for minutes. If you need a longer pause, catch the error that the call returns, read retryAfterSeconds, and schedule the work in a job queue.
What the client does by itself
The default timeout of a request is 10 minutes. The client sends one credential only, the x-api-key header, so you never hit the 401 that the API returns when a request has both Authorization and x-api-key. Generated operations resolve to { data, error, response } and do not throw, so retries happen inside the call, and what you see is the final result.
The 10 minute timeout is long because some calls, such as a sync generation wait, can hold the connection. For short reads you may want a lower value. Choose a number that suits the route, and do not copy the default into every call just because it is there.
Retry rules at a glance
The table lists the rule for each kind of response.
One row deserves attention. A 429 on a POST is retried only when the key is set, and a 429 can mean rate_limited or queue_full. For the queue case, retrying right away is rarely useful, because the queue needs room first. Wait for jobs to finish or cancel queued ones, and then retry with the same key.
| Response | Retried by the client | Condition |
|---|---|---|
| 408 or 5xx on a GET | Yes | Up to maxRetries |
| 429 on a GET | Yes | Waits for retry-after, up to 60 seconds |
| 5xx or 429 on a POST | Only with a key | An Idempotency-Key must be set |
| 400, 401, 402, 403, 404, 409 | No | Fix the request first |
The delay rule as code
The sketch below reproduces the delay rule so you can unit test your own wrapper or choose a value for maxRetries. It is not the SDK code. It follows the documented behaviour: a retry-after value wins, the wait is capped at 60 seconds, and jitter is plus or minus 20 percent.
Test the policy in two directions. Check that a transient failure on a keyed POST does get retried, and check that the same failure without a key does not. The second test protects your budget, because it proves that an unkeyed paid call never repeats.
export function delayMs(attempt: number, retryAfterSeconds?: number): number {
const base = retryAfterSeconds !== undefined
? Math.min(retryAfterSeconds, 60) * 1000
: Math.min(500 * 2 ** attempt, 8000);
const jitter = 1 + (Math.random() * 0.4 - 0.2);
return Math.round(base * jitter);
}
export function shouldRetry(status: number, method: string, hasKey: boolean): boolean {
const transient = status === 408 || status === 429 || status >= 500;
if (!transient) return false;
return method.toUpperCase() !== "POST" || hasKey;
}
console.log(shouldRetry(503, "POST", false), shouldRetry(503, "POST", true));
console.log(delayMs(0, 7) >= 5600, delayMs(0, 120) <= 72000);
Do not stack two retry loops
Pick one retry layer. If you wrap the client in your own loop and also leave maxRetries at 2, a single logical failure can turn into nine attempts. Either set maxRetries to 0 and own the policy, or keep the default and let your code handle only the final error. Do not retry a failed Format run with the same key at either layer, because the receipt of a failed run would come back unchanged.
Log the number of attempts per request. If most calls need two attempts, the problem is not the retry setting but the load, the plan limit or a degraded upstream, and a different fix is called for.
Retries and the rate limit
Retries also spend rate limit budget. Each repeat is another request in the same per-minute bucket, and a 429 names its error.details.scope, which is read or write. If a loop of status polls keeps hitting 429, lower the poll rate or move to webhooks. If submits hit it, spread them out and keep the stable key so repeats stay safe.
Sources
Related posts
More in Developers
- Sume STT silence split: it only runs after the last sentence end
Sume STT cuts at terminal punctuation first; the 0.5 s silence rule only splits the unpunctuated run after the last terminal, not earlier pauses.
- Sume sync mode: set your client timeout above 30 s, then poll
Sume sync waits at most 30 seconds. If your HTTP client gives up at 30 s too, you lose the job id. Use a 45 s timeout, then poll and reuse the Idempotency-Key.
- Sume TTS 1.0 or TTS Router: which endpoint for a voiceover?
Both bill $0.0475 per 1,000 characters. TTS 1.0 always runs sonic-3.6; the router makes you name a Sonic id. Pick by whether you need to pin the engine.
- Sume TTS 400: send transcript or transcript_source, never both
A TTS request needs exactly one of transcript or transcript_source. Both, or neither, is an error. Live-commerce Formats need the source. A validator.
Written by Sume