wave_size_hint on an empty workspace: Free 4, Pro 18, Startup 36
What generation_limits.wave_size_hint returns before any job is queued, by Sume plan, with the arithmetic and a Python check you can run.

On an empty workspace the Sume wave_size_hint is 4 on Free, 18 on Pro, 36 on Startup, and 90 on Scale. The hint is max(1, floor(queue_capacity_remaining * 0.75)), and with nothing running, queue_capacity_remaining equals the accepted job capacity of the plan. The hint sizes a submission wave. It is not a concurrency limit.
The formula, plan by plan
The Generation admission page gives the default queue capacity as max(3, concurrency_limit x 5), and accepted job capacity as concurrency_limit + queued_jobs_limit. Multiply the accepted capacity by 0.75 and round down. The table uses the plan defaults from the docs, as of 2026-10-09.
| Plan | Concurrency | Queue capacity | Accepted capacity | Arithmetic | wave_size_hint |
|---|---|---|---|---|---|
| Free | 1 | 5 | 6 | floor(6 x 0.75) = floor(4.5) | 4 |
| Pro | 4 | 20 | 24 | floor(24 x 0.75) | 18 |
| Startup | 8 | 40 | 48 | floor(48 x 0.75) | 36 |
| Scale | 20 | 100 | 120 | floor(120 x 0.75) | 90 |
Check it in Python
The same arithmetic in a few lines. It needs no network and no key, so you can use it as a unit test for your own wave planner.
import math
PLANS = {"free": (1, 5), "pro": (4, 20), "startup": (8, 40), "scale": (20, 100)}
def wave_size_hint(queue_capacity_remaining: int) -> int:
return max(1, math.floor(queue_capacity_remaining * 0.75))
for name, (concurrency, queue) in PLANS.items():
accepted = concurrency + queue
print(name, accepted, wave_size_hint(accepted))What the hint does not tell you
The hint is a snapshot of how many submits are likely to be accepted right now. It is not the number of jobs that run at once. On Free, four jobs in a wave still process one at a time, because processing concurrency is 1. Sume accepts valid jobs as queued while queue capacity remains, and workers move them to processing under the concurrency limit of the workspace.
The docs say never to show the hint as concurrency and never to use it to size in-flight work. For that, use the effective concurrency_limit minus active and queued jobs. The headroom post covers that formula.
Static tables go stale. The dashboard Concurrency tab is the source of truth for the configured processing cap, and an admin override changes concurrency_limit and the queue that follows from it. Read the fields from each submit response and treat this table as a sanity check.
When the hint reaches 1
The minimum value of the hint is 1, even with a full queue. A hint of 1 with queue_capacity_remaining: 0 does not mean that one more job fits. It means that the formula floors at 1. A full queue still returns 429 queue_full, and the client should wait for jobs to finish or cancel queued jobs. Then retry with the same Idempotency-Key.
Reading the numbers
The hint is a pacing aid, not a promise. It is computed from the queue room left at the moment of the response, so a workspace with jobs already queued gets a smaller number than the empty-workspace values above. Treat each snapshot as stale as soon as you send the next submit.
A wave of the hinted size fills about three quarters of the remaining queue and leaves room for other callers on the same workspace, such as a teammate or a scheduled job. If you are the only caller, you can send the full headroom; if you share the workspace, the hint is the safer width.
Sources
Related posts
More in Developers
- Which AI video models on Sume make their own audio track
Gemini Omni Flash 1.1 always generates audio and rejects generate_audio false; MiniMax H3 Max has native stereo audio; Seedance 2.0 reports generate_audio true.
- Which id to send Sume support: req_, job_, arun_ or agrun_
Error bodies carry req_ ids; jobs have job_; Format and Action runs arun_; Agent Completions agrun_. Which to share with support, and what never to paste.
- Which Sume wait mode fits which job: image, video, avatar, swap
Sume's async, sync, subscribe and webhook modes differ only in how you learn the outcome. A table by job length, and why 30 seconds is a request budget.
- X API video upload: chunked only, 0.5 s minimum, 20-minute cap
X's media docs require chunked upload for all videos, set a 0.5-second minimum and a 20-minute cap for non-Premium posts. Trim a clip for DMs with video_trim.
Written by Sume