Sume plan cost per concurrent slot: Pro $10, Startup $15, Scale $20

Sume Pro is $40 a month for 4 concurrent jobs, Startup $120 for 8, Scale $400 for 20. What a slot costs, how big the queue is, and what top-ups cannot buy.

5 min readSume
All posts

On Sume's public plans a concurrent generation slot costs $10 a month on Pro ($40 for 4 slots), $15 on Startup ($120 for 8) and $20 on Scale ($400 for 20). Slots are plan-only: prepaid top-ups do not raise concurrency, and generation is billed at each model's published rate. The per-slot figure is simple division of the sticker price; paid plans also carry other features (Startup adds a shared team workspace, Scale adds priority support), so it is not an itemised slot price.

The slot price rises with the tier, so the cheapest way to run many jobs at once is not always the biggest plan. The queue is what makes smaller plans workable, and the numbers below show how much it absorbs.

What does each plan give you?

Prices come from the plan catalog on Pricing; the slot and queue numbers come from Generation admission. The queue default is max(3, concurrency_limit x 5).

Free covers image generation and limited Sume Agent access. Video generation, full Sume Agent access and API, CLI and MCP access start at Pro, per the pricing FAQ.

Sume plan concurrency and queue, read 2026-10-02
PlanMonthly priceConcurrent jobsQueue (default)Accepted at oncePer slot
Free$0156n/a
Pro$4042024$10
Startup$12084048$15
Scale$40020100120$20

What does the queue do for a small plan?

Concurrency is a dispatch limit, not a submit limit. If a Pro workspace already has 4 jobs processing, Sume still accepts valid jobs as queued until accepted capacity (24 on Pro) is full. Only then does a new paid submission fail with 429 queue_full.

That means 4 slots do not stop you submitting 24 jobs. They set how many run in parallel, and therefore how long the last one waits. Docs describe video generation as taking 30 seconds to several minutes depending on model and resolution, so wait time scales with your queue depth, not with price.

What is the throughput difference?

As an assumption only, take a job that runs 3 minutes. One slot finishes 20 an hour, so the plan multiplies that.

The last column is the same because each plan's queue is five times its slots: capacity and throughput scale together. What differs is the absolute batch you can park at once, 24, 48 or 120, which is how you decide: pick the plan whose accepted capacity covers your biggest single batch.

Illustrative hourly throughput if every job takes 3 minutes (assumption, not a measured figure, read 2026-10-02)
PlanSlotsJobs per hourHours to clear a full accepted queue
Pro4800.3
Startup81600.3
Scale204000.3

What can I not buy with a top-up?

The docs are explicit that prepaid top-ups do not raise processing concurrency. A top-up funds the balance that jobs reserve against, and a submit fails with 402 insufficient_credits when the estimate cannot be reserved. Teams can fund a shared wallet with on-demand purchases from $10 to $1,000, according to the pricing FAQ.

Org workspaces have a floor of 10 concurrent jobs, and Enterprise defaults to 20 with admin overrides for contract limits. Always read the effective generation_limits.concurrency_limit on a submit response rather than trusting the static table.

How do I size a batch against my plan?

Compute new in-flight work as concurrency_limit - active - queued, capped by queue_capacity_remaining, and pace submissions to it. The wave_size_hint field, max(1, floor(queue_capacity_remaining x 0.75)), is a submission-wave hint only, never a concurrency limit.

If you hit queue_full, wait for jobs to finish or cancel queued ones, then retry with the same idempotency key.

When is a bigger plan the wrong answer?

When your batches are small. If your largest single batch is 20 jobs, Pro's accepted capacity of 24 covers it and the extra $80 a month for Startup buys parallelism you do not need.

When your batches are large and time matters, the slot count is what you are buying. The per-slot price rises from $10 to $20 as you climb, so the premium pays for headroom and not for efficiency.

Read the effective limits from a submit response before you decide. If an admin override raises your concurrency, the static table above no longer describes your workspace.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume