30-clip storyboard agent: which Sume plan accepts all jobs at once?

If an agent submits 30 separate generation jobs, only Startup (48 accepted) and above take them all at once; Pro (24) needs two waves of 18 and 12.

5 min readSume
All posts

A storyboard agent that submits 30 separate paid generation jobs is accepted in full only on Startup (48 accepted jobs), Scale or Enterprise (120). Pro accepts 24, so the last six get 429 queue_full, and Free accepts 6. Accepted capacity is the processing concurrency plus the queue, and jobs beyond it are refused rather than queued.

Processing is separate: on Pro only 4 jobs run at once, so 30 accepted jobs still finish in waves of four. Accepting a job is about the queue, not about speed.

Plan limits from the docs

Defaults from the Generation admission page, read 2026-10-09. The dashboard Concurrency tab and generation_limits in the API response are the source of truth for your workspace.

Accepted capacity = concurrency + queue; wave hint = max(1, floor(remaining x 0.75))
PlanProcessingQueueAcceptedFits 30 at once?wave_size_hint (empty queue)
Free156No4
Pro42024No18
Startup84048Yes36
Scale20100120Yes90
Enterprise20100120Yes90

What the agent should do on Pro

On Pro, with an empty queue, queue_capacity_remaining is 24 and the hint is max(1, floor(24 x 0.75)) = 18. Submit 18, wait until they drain, then submit the remaining 12. The hint is only a guide for submission waves; it is not a concurrency limit, and the docs say not to use it to size in-flight work.

To size work in flight, use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs) capped by queue_capacity_remaining, refreshed from a live generation_limits snapshot. Refresh before every wave.

import math

def wave_hint(remaining):
    return max(1, math.floor(remaining * 0.75))

plans = {"free": 6, "pro": 24, "startup": 48, "scale": 120}
for name, accepted in plans.items():
    print(name, accepted, wave_hint(accepted), accepted >= 30)

Retry rules when the queue is full

429 queue_full is retryable: wait for jobs to finish or cancel queued ones, then retry with the same idempotency key. It is different from 429 rate_limited, which is request volume, and from 402 insufficient_credits, which means the balance cannot be reserved. For the clip prices themselves, use the public catalog; a 5-second Wan 3.0 720p clip is 5 x $0.125 = $0.625, so 30 of them are $18.75 before any retries.

What it means for the budget

Capacity and money are separate checks. A plan that accepts 30 jobs still needs the balance to reserve each estimate, or the submit fails with 402 insufficient_credits. For 30 five-second Wan 3.0 720p clips the list total is 30 x $0.625 = $18.75. At the Seedance 2 720p rate of $0.378 per second the same 30 clips are 30 x 5 x $0.378 = $56.70. Preview with dry_run or generation_admission_preview before the first wave, and read generation_limits from the response to see how much room is left.

Also watch for an unusual case: org workspaces have a floor of 10 processing seats and Enterprise can be raised by admin override, so the effective fields in the response beat this table.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume