Size a Sume submit wave: limit 100, 30 processing, 10 queued gives 60
How many new jobs can a Sume workspace take? max(0, limit - active - queued), capped by queue_capacity_remaining. Worked numbers and a TypeScript helper.

The in-flight budget for new Sume jobs is max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped at queue_capacity_remaining. With a limit of 100, 30 jobs processing, and 10 queued, that is 100 - 30 - 10 = 60. The wave_size_hint field is not that number, and Sume says never to use it to size in-flight work.
Generation concurrency on Sume is a dispatch limit, not a submit limit: a workspace at its limit still accepts valid jobs as queued while queue capacity remains. The budget above keeps your own waves inside the processing cap, so your queue does not fill with work you did not need to submit yet.
Two snapshots
The idle row uses the example in the Generation admission page. The busy row applies the page's definitions of queue_capacity_remaining (remaining queued budget plus idle processing seats) and wave_size_hint (max(1, floor(queue_capacity_remaining * 0.75))). The arithmetic is mine.
| Snapshot | queue_capacity_remaining | wave_size_hint | In-flight budget |
|---|---|---|---|
| Idle: 0 processing, 0 queued | 600 | 450 | min(100, 600) = 100 |
| Busy: 30 processing, 10 queued | (500 - 10) + (100 - 30) = 560 | floor(560 x 0.75) = 420 | min(100 - 30 - 10, 560) = 60 |
The helper
Refresh the snapshot from the generation_limits object on each submit response, and count every job you submit against the budget until the next snapshot. The helper returns how many to submit now.
type GenerationLimits = {
concurrency_limit: number;
active_generation_jobs: number;
queued_generation_jobs: number;
queue_capacity_remaining: number;
};
export function newInFlightBudget(l: GenerationLimits): number {
const headroom = Math.max(
0,
l.concurrency_limit - l.active_generation_jobs - l.queued_generation_jobs,
);
return Math.min(headroom, l.queue_capacity_remaining);
}
// { concurrency_limit: 100, active: 30, queued: 10, remaining: 560 } -> 60When the budget is zero
At zero headroom, wait and refresh before you submit more. If counts are unavailable, refresh before you pick a width. The snapshot can change right after the response, since workers claim jobs and other clients submit work.
The hard stop is 429 queue_full, which means no accepted capacity is left. The docs say to wait for jobs to finish, cancel queued jobs you no longer need, and retry with the same idempotency key. Do not show wave_size_hint as concurrency in a dashboard, and do not read the old plan_concurrency_limit as the effective limit when limit_source is admin_override.
Why not use wave_size_hint
The hint is a suggested submit size, computed from remaining queue capacity. It is not the dispatch limit. With the busy snapshot above it reads 420, while the processing cap leaves room for only 60 new running jobs. If you submit 420 jobs, 360 of them sit queued behind the cap, and they hold queue capacity other work could have used.
That is also why the helper reads concurrency_limit. The same field can come from your plan or from an admin override, and the response names which with limit_source. Use the effective value that the response gives you, not a plan table you copied last month.
A simple rule: submit up to the budget, wait for terminal jobs through webhooks or polling, then recompute. Each submit response carries a fresh generation_limits, so there is no extra read to make.
Sources
Related posts
More in Developers
- Smoke-test 1:4, 4:1, 1:8 and 8:1 on Nano Banana 2.1 for $0.30
Four calls at 0.5K, one per new ratio, cost $0.30 on Sume ($0.80 at 4K). A script that prints status and cost per ratio, and what a 400 or 202 means.
- Sora, Veo and Omni preview shutdown dates in one table (Oct 2026)
OpenAI removed the Sora Videos API on Sept 24, 2026. Google ends Veo 3.1 and Omni preview ids Oct 22. One dated table, and how to read Sume's live list.
- Make an SRT file from Sume STT word timings in Python (7 words a cue)
Sume STT always returns words[] with word, start and end in seconds. This Python function turns the list into an SRT file, seven words a cue.
- Stable idempotency keys for pause-cut trims: hash source, start, end
Derive each trim's Idempotency-Key from a hash of source, start, end and precision. A retried batch then returns the same jobs and bills 8 trims once.
Written by Sume