37 video jobs submitted at once: which Sume plan accepts them all

Free accepts 6 of 37 video jobs, Pro 24, Startup all 37 with 11 spare slots, Scale all 37. Concurrency, queue and accepted-capacity table from the Sume docs.

5 min readSume
All posts

Startup is the smallest Sume plan that accepts 37 simultaneous paid video jobs: its accepted capacity is 8 processing + 40 queued = 48. Pro accepts 24, so 13 of the 37 would return 429 queue_full, and Free accepts 6, so 31 would. Scale accepts 120. This holds for any paid generation job, video included, because the limits belong to the workspace, not to the model.

Accepted capacity by plan

Numbers are from the Sume generation admission page, read 2026-10-09. Accepted capacity is concurrency plus queue capacity. Prepaid top-ups do not raise processing concurrency; it is plan-only, though admin overrides can raise it, so check your workspace's effective concurrency_limit.

37 simultaneous submits against each plan, per the admission docs read 2026-10-09
PlanProcessingQueueAccepted capacityAccepted of 37Rejected
Free156631
Pro420242413
Startup84048370
Scale20100120370

The money side

Each accepted job holds its price at submit. If the 37 jobs are 5-second Wan 3.0 clips at 720p, each is 5 x $0.125 = $0.625, so a full set is 37 x $0.625 = $23.125. On Pro, the 24 accepted jobs hold $15.000 and the 13 rejected ones hold nothing. A workspace balance below the total returns 402 insufficient_credits for the first submit that the balance cannot cover, which is a separate check from the queue.

Processing speed versus acceptance

Acceptance and speed are different things. Startup accepts all 37 jobs at once, but it processes only 8 at a time, so jobs 9 to 37 wait as queued. Scale accepts 37 and processes 20 at a time. The plan decides how fast the batch drains, and the queue capacity decides only whether a job is accepted.

If time to finish matters, divide: 37 jobs on 8 slots take five rounds (8 + 8 + 8 + 8 + 5), and on 20 slots two rounds (20 + 17), assuming equal job length. Real generation times vary by model and resolution, so treat the round count as a bound on the structure, not a time estimate.

Use the effective limits in your own workspace as the source of truth. Org workspaces have a floor of 10 for concurrency and Enterprise starts at 20 with admin overrides for contract limits, so the table is a floor for planning and not a guarantee for every account. The API reports the cap as generation_limits.concurrency_limit.

Submitting in waves instead

If you stay on Pro, split 37 into a wave of 24 and a wave of 13. Keep the job ids, poll with backoff or use webhooks, and send the second wave when the first has drained enough slots. Each slot frees as a job reaches completed, failed or canceled. A retry with the same Idempotency-Key returns the original job, so a lost response does not double-bill.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume