Can polling Sume job status hit the rate limit? Read budget math
Reads and writes have separate Sume budgets. Worked numbers for polling every accepted job per plan, plus a delay function that reads ratelimit headers.

Polling job status will not exhaust a Sume key's rate limit if you poll at the suggested 2-second pace, because status reads draw from a read budget that is forty times the write budget. Even polling every accepted job on a plan once per second stays well under the per-minute read number. The limit you are more likely to meet is queue_full, which is generation capacity, not request rate.
The arithmetic below uses only figures from the Sume docs. A job counts toward polling while it is queued or processing, and the most jobs a workspace can hold in those states is its accepted job capacity (concurrency plus queue).
Worked numbers
Each cell is accepted jobs times polls per minute. Writes per minute are 120, 300, 600 and 1200 for the same plans, so submits and cancels are the budget to watch, not status reads. Enterprise is contract-specific, and until a number is provisioned resolves to the Scale row.
| Plan | Accepted jobs | Reads/min at 2s polls | Reads/min at 1s polls | Read budget/min |
|---|---|---|---|---|
| Free | 6 | 180 | 360 | 4800 |
| Pro | 24 | 720 | 1440 | 12000 |
| Startup | 48 | 1440 | 2880 | 24000 |
| Scale | 120 | 3600 | 7200 | 48000 |
What still goes wrong
The docs also say to read ratelimit-remaining rather than counting requests yourself, since the headers describe whichever budget the current request spent from.
- One key shared by many services: the budget is per API key, so a second service polling with the same key spends the same read bucket.
- Polling without
next_poll_after_seconds: hammering a status URL at 100 ms is allowed by the budget on paper but wastes the budget for everything else. - Reading the wrong signal: a 429 names its scope in
error.details.scope,readorwrite. A write 429 on submit does not mean your poller is too fast.
A delay function that reads the headers
It returns how long to sleep before the next status read. A missing header is treated as unknown, not zero, because Number(null) is 0 and would look like an empty budget.
// Read the budget the response just spent from. Headers can be absent, and
// Number(null) is 0, so test for null before converting.
const num = (res, name) => {
const v = res.headers.get(name);
return v === null ? null : Number(v);
};
export function nextPollDelayMs(res, hintSeconds = 2) {
const base = Math.max(2, hintSeconds) * 1000;
const limit = num(res, "ratelimit-limit");
const left = num(res, "ratelimit-remaining");
const reset = num(res, "ratelimit-reset");
if (res.status === 429) {
const wait = num(res, "retry-after") ?? reset ?? 5;
return wait * 1000;
}
if (limit && left !== null && reset !== null && left < limit * 0.1) {
return Math.max(base, reset * 1000); // under 10% left: wait for the window
}
return base;
}Limits
I tested the function with hand-made responses: no headers gave 2000 ms, healthy headers gave the hint, 900 of 12000 remaining waited the 30-second reset, and a 429 used retry-after. Those are not live headers. The read multiple is a deployment setting, and the docs say ratelimit-limit on the response is the authority for the deployment you call, so the table is the shipped default and not a guarantee. Unauthenticated requests get a much smaller read bucket, so never poll without a key.
Sources
Related posts
More in Developers
- Cancel queued Sume jobs on SIGTERM during a deploy
On shutdown, cancel jobs you no longer need before they start and leave started ones alone. A 21-line Node handler using POST /v1/jobs/{id}/cancel.
- Check a video request against /v1/videos/models before you submit
Duration, resolution, size and seed errors cost a round trip. A short Python validator reads the model catalog and refuses a bad request locally first.
- Check a transparent GPT Image 2.5 PNG for real alpha in Python
A transparent GPT Image 2.5 result can still look opaque. Ask for background transparent as PNG, then check the alpha channel in Python: a 20-line script.
- Claude structured outputs drop minimum and maxLength; Sume keeps them
Anthropic lists minimum, maxLength and recursion as unsupported in structured outputs. Sume's output_schema accepts the first two; here is what differs.
Written by Sume