Size a batch from generation_limits so no clip hits queue_full

Read accepted_generation_jobs_limit from a Sume submit response and slice your clips. A 50-clip batch leaves 2 for a later wave on Startup and 26 on Pro.

5 min readSume
All posts

Before a bulk submit, read generation_limits from any generation response and cap the number of jobs you have in flight at accepted_generation_jobs_limit. That value is the processing concurrency plus the queue, for example 4 + 20 = 24 on Pro. Past it, new paid submissions fail with 429 queue_full. With cheaper clips arriving this autumn, such as Veo 3.1 Lite at $0.05 per second at 720p on Google's pricing page, people queue bigger batches, and this limit is the one they hit.

Concurrency is not the limit you submit against

Sume treats concurrency as a dispatch limit, not a submit limit. A workspace at its concurrency limit still accepts new valid jobs as queued while queue capacity remains. fal's queue docs describe the opposite posture: requests are never dropped and runners scale to demand. Sume bounds the queue, so your client should too.

What a 50-clip batch does by plan (read 2026-10-05)
PlanAccepted job capacityAccepted at onceLeft for a later wave
Free6644
Pro242426
Startup48482
Scale120500

Slice the list

The function below takes how many jobs are still active and returns the next slice. The limits dictionary mirrors the fields in a submit response.

def next_wave(pending: list, limits: dict, active: int):
    room = limits["accepted_generation_jobs_limit"] - active
    room = max(room, 0)
    return pending[:room], pending[room:]

limits = {
    "concurrency_limit": 4,
    "queued_jobs_limit": 20,
    "accepted_generation_jobs_limit": 24,
}
clips = [f"clip-{i}" for i in range(50)]
submit, rest = next_wave(clips, limits, active=0)
print(len(submit), len(rest))  # 24 26
submit, rest = next_wave(rest, limits, active=20)
print(len(submit), len(rest))  # 4 22

Rules that keep it safe

Count a job as active from the moment it is queued or processing, and stop counting on a terminal status. If you do hit queue_full, nothing is lost: wait for a job to finish or cancel queued jobs, then retry with the same Idempotency-Key. Prefer the effective concurrency_limit in the response over a static table, because an admin override can change it.

  • Use one Idempotency-Key per clip so a retry returns the original job.
  • Poll status with backoff instead of resubmitting.
  • A Format bulk run does the windowing for you server-side, up to 100 items and a concurrency of 1 to 16.

Edge cases worth a test

Test three cases before a nightly run. An empty pending list should submit nothing. A limit response that is missing from a failed submit should fall back to your last known limits rather than to zero. And a restart in the middle of a batch should rebuild active from GET /v1/jobs, filtered to queued and processing, not from memory.

Sume exposes generation_limits on generation submit responses when it can compute them, with limit_source telling you whether the number comes from the plan or from an admin override. Use concurrency_limit, the effective value, not plan_concurrency_limit, when you plan waves.

Finally, cap the cost of a bug. A loop that forgets to count active jobs will hit queue_full quickly, which is the safe failure. The unsafe one is an un-keyed retry, so keep one idempotency key per clip.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume