Size a batch from generation_limits so no clip hits queue_full
Read accepted_generation_jobs_limit from a Sume submit response and slice your clips. A 50-clip batch leaves 2 for a later wave on Startup and 26 on Pro.

Before a bulk submit, read generation_limits from any generation response and cap the number of jobs you have in flight at accepted_generation_jobs_limit. That value is the processing concurrency plus the queue, for example 4 + 20 = 24 on Pro. Past it, new paid submissions fail with 429 queue_full. With cheaper clips arriving this autumn, such as Veo 3.1 Lite at $0.05 per second at 720p on Google's pricing page, people queue bigger batches, and this limit is the one they hit.
Concurrency is not the limit you submit against
Sume treats concurrency as a dispatch limit, not a submit limit. A workspace at its concurrency limit still accepts new valid jobs as queued while queue capacity remains. fal's queue docs describe the opposite posture: requests are never dropped and runners scale to demand. Sume bounds the queue, so your client should too.
| Plan | Accepted job capacity | Accepted at once | Left for a later wave |
|---|---|---|---|
| Free | 6 | 6 | 44 |
| Pro | 24 | 24 | 26 |
| Startup | 48 | 48 | 2 |
| Scale | 120 | 50 | 0 |
Slice the list
The function below takes how many jobs are still active and returns the next slice. The limits dictionary mirrors the fields in a submit response.
def next_wave(pending: list, limits: dict, active: int):
room = limits["accepted_generation_jobs_limit"] - active
room = max(room, 0)
return pending[:room], pending[room:]
limits = {
"concurrency_limit": 4,
"queued_jobs_limit": 20,
"accepted_generation_jobs_limit": 24,
}
clips = [f"clip-{i}" for i in range(50)]
submit, rest = next_wave(clips, limits, active=0)
print(len(submit), len(rest)) # 24 26
submit, rest = next_wave(rest, limits, active=20)
print(len(submit), len(rest)) # 4 22Rules that keep it safe
Count a job as active from the moment it is queued or processing, and stop counting on a terminal status. If you do hit queue_full, nothing is lost: wait for a job to finish or cancel queued jobs, then retry with the same Idempotency-Key. Prefer the effective concurrency_limit in the response over a static table, because an admin override can change it.
- Use one
Idempotency-Keyper clip so a retry returns the original job. - Poll status with backoff instead of resubmitting.
- A Format bulk run does the windowing for you server-side, up to 100 items and a concurrency of 1 to 16.
Edge cases worth a test
Test three cases before a nightly run. An empty pending list should submit nothing. A limit response that is missing from a failed submit should fall back to your last known limits rather than to zero. And a restart in the middle of a batch should rebuild active from GET /v1/jobs, filtered to queued and processing, not from memory.
Sume exposes generation_limits on generation submit responses when it can compute them, with limit_source telling you whether the number comes from the plan or from an admin override. Use concurrency_limit, the effective value, not plan_concurrency_limit, when you plan waves.
Finally, cap the cost of a bug. A loop that forgets to count active jobs will hit queue_full quickly, which is the safe failure. The unsafe one is an un-keyed retry, so keep one idempotency key per clip.
Sources
Related posts
More in Developers
- Voice model updated in place with no API change: how to detect it
Nova 2 Sonic was refreshed in place in May with no API change. If a vendor can change your voice silently, log the model id and a canary clip. Sume code inside.
- spend_approval_queue_full 429: clear pending approvals, do not retry
A thread with too many pending spend approvals gets 429 spend_approval_queue_full. Resolve the pending ones first; a 503 store_misconfigured is for support.
- spend_confirmation_required 402 on Sume: why retrying will not help
A 402 spend_confirmation_required means a person must approve the spend first. It is not a balance error: retryable is false and next_action is fix_input.
- Split a 10-minute TikTok into parts for 3 and 5-minute accounts
TikTok's API allows up to 10 minutes, but an account may be limited to 3 or 5. Split a 600-second video into 4 or 2 parts with Sume trim at $0.02 a job.
Written by Sume