Sume SDK 429 retry-after is capped at 60 seconds: what follows

createSumeClient waits at most 60 seconds per retry, with 20 percent jitter and 2 retries by default. What that means for a long retry-after and a fix.

4 min readSume
All posts

createSumeClient honours a retry-after header, but it never sleeps longer than 60 seconds for one retry, it adds up to 20 percent jitter either way, and by default it retries only twice. If a 429 ever asks for more than a minute, the client will come back early, probably get a second 429, and after the second retry return the failure to you. The fix is to read retryAfterSeconds off the final error and wait that long yourself.

What exactly does the client do between attempts?

The docs state the headline: createSumeClient retries 408, 429, 5xx and transport failures, 2 retries by default, with exponential backoff and jitter, honouring retry-after. The package source fills in the arithmetic.

Retry timing in createSumeClient. Source: TypeScript SDK and Waiting for runs and jobs at docs.sume.com, read 2026-10-03; the cap and jitter come from the SDK source in the Sume repository.
CaseBase waitAfter 20% jitter
retry-after: 55 s4 to 6 s
retry-after: 3030 s24 to 36 s
retry-after: 12060 s (capped)48 to 72 s
No header, first retry0.5 s0.4 to 0.6 s
No header, second retry1 s0.8 to 1.2 s

Does this change when I should worry?

Only for unusually long windows. The per-key budget is per minute, set by the plan: writes are 120 a minute on Free, 300 on Pro, 600 on Startup and 1200 on Scale, and reads get forty times the write number. A window measured in seconds fits inside the cap. Read ratelimit-reset, which says seconds until the window resets, if you want to know rather than guess. The cap matters when a response names a longer wait than the client will give.

Retries are also limited by request type. A GET is always replayed. A POST is replayed only when it carries an Idempotency-Key, so an unkeyed run create returns its first 429 straight to you.

A wrapper that waits as long as the server asked

Turn the built-in retries down and own the loop. This keeps each submit on one Idempotency-Key, so waiting longer never starts a second paid run.

import { createSumeClient, toSumeApiError } from "@sume-com/sdk";

const client = createSumeClient({
  apiKey: process.env.SUME_API_KEY!,
  maxRetries: 0, // we handle 429 ourselves
});

export async function withRetryAfter(call, tries = 4) {
  for (let attempt = 1; ; attempt += 1) {
    const res = await call(client);
    if (!res.error) return res.data;
    const err = toSumeApiError(res.response?.status, res.error, "request failed");
    if (err.status !== 429 || attempt >= tries) throw err;
    const wait = err.retryAfterSeconds ?? 5;
    await new Promise((r) => setTimeout(r, wait * 1000));
  }
}

What if the 429 is queue_full, not rate_limited?

There are two 429 codes. rate_limited means too many requests in the current window. queue_full means the workspace's generation concurrency plus queue capacity is full, and request pacing does not change that. Branch on err.code before you sleep: pacing fixes the first and only finished generations fix the second.

How do I see my headroom before I hit a 429?

Every response can carry ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after. The first three describe the budget the request spent from, writes and reads being separate, and retry-after is sent on a 429. Pacing on the headers beats counting requests yourself, because the same code can sit well inside the limit on a Scale key and over it on a Free one.

A practical rule: if ratelimit-remaining drops under a small reserve, wait for ratelimit-reset seconds before the next burst. Do this on the submit path, where an extra wait costs nothing, and leave status polling on the SDK's own jittered intervals, which exist so that several clients started together do not stay in phase.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume