Do more credits make Sume jobs faster? Concurrency is plan-only

No. Sume's processing concurrency is set by plan, not by balance: Free 1, Pro 4, Startup 8, Scale 20. Top-ups only raise what you can reserve.

4 min readSume
All posts

No. Adding credits to a Sume workspace does not make jobs run faster, because generation concurrency is plan-only. The docs say it plainly: prepaid top-ups do not raise the processing concurrency limit. Your balance decides how much work you can reserve, and your plan decides how many paid jobs process at the same time.

This matters when a big batch is slow and the instinct is to add money. Money changes whether a submit is accepted. It does not change how many jobs run in parallel.

What does balance control, and what does the plan control?

The generation admission doc separates four controls that are easy to confuse: generation concurrency, queue capacity, submit rate limits, and balance with reservation. Only the last one is about money. When the balance is short, the submit fails with 402 insufficient_credits before provider work starts.

Concurrency is a dispatch limit, not a submit limit. A workspace at its limit can still accept more valid jobs as queued while queue capacity remains, and workers move them to processing as slots open. Queue capacity defaults to max(3, concurrency_limit x 5).

What does each plan give?

The table is the default map from the admission doc. The dashboard Concurrency tab is the source of truth for a workspace, exposed as generation_limits.concurrency_limit; org workspaces have a floor of 10, Enterprise defaults to 20 and uses admin overrides for higher contract limits.

The last column is arithmetic: how many jobs each processing slot works through in a row if you fill the accepted capacity. It is a floor on how many turns a slot takes, not a time estimate.

Default processing concurrency and queue capacity by plan (Generation admission doc, read 2026-10-03); last column is accepted jobs / concurrency, rounded up.
PlanProcessing concurrencyQueue capacityAccepted jobsJobs per slot when full
Free1566
Pro420246
Startup840486
Scale201001206
Enterprise201001206

How long does a queue take to drain?

Take 24 identical jobs. On Pro, all 24 are accepted and each of the 4 slots handles about 6 jobs in turn. On Startup, 8 slots handle 3 each. On Scale, 20 slots handle at most 2. The finishing time is roughly the job length times the turns, so the plan changes the wall-clock time and the balance does not.

On Free, concurrency is 1 and accepted capacity is 6, so a 24-job batch cannot be accepted all at once. The seventh job returns 429 queue_full until one job finishes. Paying more does not widen that queue. Moving up a plan does.

So what do I do with a bigger balance?

Use it to keep every submit above the reserve: a bulk run needs balance for every accepted job at once, since each one holds its estimate. A $200 balance on Pro does not run jobs sooner than $50, but it avoids a 402 partway through a large queue.

To go faster, change the plan or ask about an override. Admin overrides can raise the effective concurrency_limit (shown as limit_source: admin_override), and the field you should always read is the effective one, not the static table.

Top-ups happen in the dashboard under Billing & subscription, which can start a Stripe-backed manual top-up when billing is configured. The public API exposes GET /v1/balance and GET /v1/usage, not a top-up endpoint. Rates are on the API pricing page.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume