How many 30-second videos can a Sume Pro key run at once? Wave math

Pro runs 4 generations at once and queues 20 more, so 24 accepted jobs. Use the plan table and a Python function to plan a 50-clip batch in waves.

4 min readSume
All posts

How many long videos can a Pro key have in flight?

Four generations run at once and twenty more can wait in the queue, so a Pro key holds 24 accepted generation jobs. A 25th submit gets a 429 queue_full. The numbers come from Sume's generation admission docs and are set by the workspace plan.

This matters most for long clips. Seedance 2.5 and Wan 3.0 both run up to 30 seconds per request (Seed, Alibaba), and a batch of them fills a queue quickly. Sume does not promise a wall-clock time for a generation, so size waves by capacity and poll, not by a guess at minutes.

The first number is the one people forget. Concurrency is how many generations the provider side is actively working on for your key, and the queue is where accepted jobs wait their turn. Both count toward the accepted total, so a queued job already uses capacity even though nothing is running for it yet.

Capacity by plan

Accepted capacity is concurrency plus queue. Enterprise matches Scale.

Sume generation capacity by plan (docs.sume.com, read 2026-10-06)
PlanRunning at onceQueueAccepted in total
Free156
Pro42024
Startup84048
Scale20100120

Plan a batch in Python

The API also returns a generation_limits snapshot with queue_capacity_remaining and a wave_size_hint of max(1, floor(queue_capacity_remaining * 0.75)). It is a hint for how many to submit next, not a concurrency number. The function below does the same arithmetic offline.

PLANS = {"free": (1, 5), "pro": (4, 20), "startup": (8, 40), "scale": (20, 100)}

def accepted(plan: str) -> int:
    running, queued = PLANS[plan]
    return running + queued

def waves(clips: int, plan: str) -> int:
    return -(-clips // accepted(plan))

def wave_size_hint(queue_capacity_remaining: int) -> int:
    return max(1, queue_capacity_remaining * 3 // 4)

for plan in PLANS:
    print(plan, accepted(plan), waves(50, plan))
print(wave_size_hint(20), wave_size_hint(1))

Submit the next wave when slots free up

Count a job as using a slot until it is terminal. Submit the next group when your last read of generation_limits shows room, or when a webhook tells you a job finished. Use a fresh Idempotency-Key per clip, and reuse it only to retry that same clip.

A 429 queue_full is not an error to hammer. Sume released the reservation, so nothing was charged: wait, then submit again with the same key. rate_limited is a different 429: it is a per-minute request budget and carries a retry-after header you should honour.

Pro also has the shortest runway of the paid plans for a big launch day. If you expect bursts above 24 clips, either spread them across waves as above or move to a plan whose accepted total fits the burst, and keep your retry code on the same Idempotency-Key so a repeated submit is a replay and never a second charge.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume