Sume SDK backoff: 500 ms to 8 s, retry-after capped at 60 s
The Sume TypeScript client waits min(500 x 2^attempt, 8000) ms, honors retry-after up to 60 s, and adds 20 percent jitter. The full table by attempt.

Short answer
With no retry-after header, the @sume-com/sdk client waits min(500 x 2^attempt, 8000) milliseconds before a retry, with 20 percent jitter either way. With a numeric retry-after, it waits that many seconds, capped at 60, with the same jitter. By default it retries twice, so in practice the waits are about 0.5 s and then about 1 s.
The formula in a table
The base delay doubles from 500 ms and stops growing at 8 seconds. Jitter multiplies the base by a random factor between 0.8 and 1.2, so that many clients that failed together do not all return at the same instant. The table lists the base and the jittered range for each attempt number, counting from 0.
| Attempt | Base delay | Jittered range |
|---|---|---|
| 0 (first retry) | 500 ms | 400 to 600 ms |
| 1 | 1000 ms | 800 to 1200 ms |
| 2 | 2000 ms | 1600 to 2400 ms |
| 3 | 4000 ms | 3200 to 4800 ms |
| 4 and later | 8000 ms | 6400 to 9600 ms |
When the server names the window
The API sends retry-after on every 429, and the client honors it instead of guessing. The value is read as a number of seconds, multiplied by 1000 and capped at 60 seconds, then jittered like any other delay, so the longest possible single wait is 72 seconds. A retry-after of zero, a negative number or a value that is not a plain number is ignored and the exponential schedule applies. That means an HTTP-date form of the header would not be honored by this client.
A 429 comes in two kinds. rate_limited means the key exceeded its per-minute read or write budget, and error.details.scope says which. queue_full means generation concurrency plus queue capacity is full. Both carry retry-after, and waiting that long is the right move for both.
What gets retried, and what does not
Only 408, 429 and 5xx statuses and transport failures are retried, and only while attempts remain. A caller abort is never retried: if you pass your own signal and abort it, the client rethrows. The per-request timeout is separate; it defaults to 10 minutes and a value of 0 turns it off.
POST needs an Idempotency-Key to be replayed. Without it the first response, even a 503, is returned as is, because a replay of a paid request without a key could bill twice. See POST retries only with a key.
Tuning it
Raise maxRetries if your callers can wait; every extra attempt adds a delay that reaches the 8 second ceiling quickly, so total wait grows linearly after the fourth retry. Lower it, or set it to 0, when you already wrap the client in a queue with its own policy; stacking two retry layers multiplies attempts. Remember that read and write budgets are separate and per key, so a burst of retried writes can drain the write budget without touching reads.
None of this replaces polling logic. Job status reads carry their own next_poll_after_seconds hint, and a retry delay on a failed read is a different thing from the interval between successful polls.
Sources
Related posts
More in Developers
- Unit test Sume SDK retries with a fake fetch
createSumeClient takes a fetch option for tests. Feed it a 503 then a 200 to prove the SDK retries a GET and sends one x-api-key, with no network.
- Cursor mcp.json ${env:NAME} for Sume's API key: no secret in the repo
Cursor's mcp.json interpolates ${env:NAME} in headers. Keep Sume's API key in an environment variable, send one credential, and know the fixed OAuth redirects.
- Cut a voiceover into sentence clips with TTS segmentation
Sume TTS returns gapless sentence segments, cutting 70 ms after each last word by default. Per-segment audio needs wav or raw; mp3 returns timings only.
- Daily spend cap in Python from a monthly Sume budget
Turn a $300 monthly budget into a $10 daily cap and enforce it with a short Python check against GET /v1/balance. Daily-cap table for $100 to $1,000.
Written by Sume