Where to read your video concurrency: dashboard tab, not a table

Sume lists plan concurrency in a table but says the dashboard Concurrency tab and generation_limits.concurrency_limit are the source of truth. How to read them.

5 min readSume
All posts

If you plan a batch of video jobs, read your concurrency from the workspace, not from a pricing table. The generation admission page prints a plan table, then says the dashboard Concurrency tab is the source of truth for the configured processing cap, exposed as generation_limits.concurrency_limit, and that you should prefer the effective field over the static table.

The reason is practical. Org workspaces have a floor of 10, Enterprise defaults to 20 and uses admin overrides for higher contract limits, and an admin override changes the effective value (limit_source: admin_override). A hard-coded 4 for Pro will undercount an org workspace and overcount one that was lowered.

The static table, as of 2026-10-03

Use it for planning conversations, not for code. Accepted job capacity is the processing limit plus the queue limit, and queue capacity defaults to the larger of 3 and five times the concurrency limit.

Plan concurrency and queue capacity on the admission page (read 2026-10-03)
PlanProcessing concurrencyQueue capacity (default)Accepted job capacity
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

What is plan-only

Generation concurrency is plan-only: prepaid top-ups do not raise the processing limit. Admin overrides can raise it. That means a bigger balance buys more reservations, not more simultaneous renders. If a team assumes otherwise, their queue fills while their balance does not move.

How to read it at runtime

These four checks take a minute and replace a guess with the numbers Sume is actually enforcing for the workspace.

  • Open the dashboard Concurrency tab for the configured processing cap.
  • Read generation_limits on generation submit responses; the docs say to use them for conservative queue decisions.
  • Call GET /v1/balance before a bulk run so the reservation and the queue are both checked.
  • Do not substitute plan_concurrency_limit or purchased_concurrency_limit for the effective concurrency_limit.

Why this matters for video

Video jobs run longer than image jobs, so a full processing slot is occupied for longer. Sume accepts valid jobs as queued while queue capacity remains, so a batch on a one-slot plan does not fail; it runs one clip at a time. If the queue itself fills, new paid submits return 429 queue_full. Size your batch to the accepted capacity, or release work in waves.

A planning example

Suppose a team has 40 clips to render and a workspace with a processing limit of 4 and a queue limit of 20. Accepted capacity is 24, so the first wave is at most 24 submits, four rendering and twenty waiting. The remaining 16 wait on your side until jobs finish. Release them as slots open and you never meet queue_full.

On a plan with a limit of 8 and a queue of 40, accepted capacity is 48, so all 40 clips are accepted at once and eight render together. The balance still has to cover 40 reservations, which is the number to check first.

Limits

The table is a snapshot, the docs say so, and they do not give a typical clip time for any plan, so this post gives no estimate of how long a batch takes.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume