Sume org workspaces run 10 generations at once; Free runs 1

Concurrency on Sume is plan-based: Free 1, Pro 4, Startup 8, Scale 20, with a floor of 10 for org workspaces. Prepaid top-ups do not raise it.

4 min readSume
All posts

Short answer

Processing concurrency on Sume depends on the plan, not on how much money is in the wallet. The generation admission docs list Free at 1 processing job, Pro 4, Startup 8, Scale 20 and Enterprise 20. Org workspaces have a floor of 10, and Enterprise has a default of 20 with admin overrides for higher contract limits (Generation admission).

The numbers

Queue capacity defaults to max(3, concurrency_limit x 5), and accepted capacity is processing plus queue.

Default concurrency and queue by plan (Sume docs, read 2026-10-05)
PlanProcessingQueueAccepted
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

Top-ups do not change it

The docs are explicit: generation concurrency is plan-only, and prepaid top-ups do not increase it. An admin override can raise the effective concurrency_limit, and the response then reports limit_source: admin_override. The dashboard Concurrency tab is the source of truth, and the API shows the same value as generation_limits.concurrency_limit. Always read the live field rather than this table.

What this means for a batch

A queued job is not an error. With a limit of 1, you can submit six valid jobs and Sume accepts all of them as queued; they move to processing one at a time. A seventh returns 429 queue_full. The reserve for every accepted job is held in the meantime, so budget the wallet for all accepted jobs, not only the running one.

Checklist

  • Read generation_limits from a submit response and use concurrency_limit, not wave_size_hint, to size in-flight work.
  • Poll with backoff; read and status endpoints have their own rate limits.
  • Org workspace or not, a 429 queue_full is solved by waiting or cancelling queued jobs, not by topping up.

What a larger limit means in time

Concurrency translates directly into wall-clock time. If a clip takes about two minutes to render, a batch of 40 on Free takes about 80 minutes of serial work, on Pro (4 at a time) about 20 minutes, and with 10 slots about 8 minutes. The provider time is not documented per model, so measure one job and scale. Polling at a moderate interval is recommended; the video docs suggest around 30 seconds.

Queue capacity matters too. A batch larger than the accepted capacity cannot be submitted at once. For 40 jobs on Pro (24 accepted), you submit in waves and wait for jobs to finish.

Do not read wave_size_hint as the concurrency. It is max(1, floor(queue_capacity_remaining x 0.75)), and the docs state never to show it as concurrency. Also do not use the pre-override plan limit when an admin override applies; read concurrency_limit and limit_source instead.

Two practical points follow from the floor. First, a small team that signs up as an organization gets room for ten videos at once without choosing a higher plan, which changes how you schedule a campaign batch. Second, the floor is a minimum, not a replacement: a plan with a higher limit still uses that limit, so check the field the API returns and not an assumption from the plan name.

For capacity planning, compute queue room from the limit. The queue is the larger of 3 and five times the concurrency, so a limit of 10 leaves room to queue 50 more jobs behind the ten running ones, and the reserve for every accepted job is held in the wallet.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume