Sume plans: Pro $40, Startup $120, Scale $400 and what they limit

What each Sume plan sets: monthly price, concurrent jobs, queue size and API write budget. Usage is billed at each model's rate, not by the plan.

5 min readSume
All posts

Sume's self-serve plans are Pro at $40 a month, Startup at $120 and Scale at $400, with a Free plan at $0 and Enterprise by contact. What a plan changes is access and limits: how many generation jobs run at once, how many can wait in the queue, and how many API writes you can make per minute. Generation itself is billed per use at each model's published rate from your USD balance, so the plan price is not a rate card.

If you are comparing against credit-bucket tools, that is the main difference to hold on to: Sume's pricing page describes the plan grid as product access and capabilities and says usage is billed at each model's published rate. The tables below list the limits from Sume's docs and plan catalog, and then show how to work out which plan you need from your own job pattern.

What does each plan cost and include?

Monthly prices come from the plan catalog that the public pricing page and the dashboard both read. The feature lines are the ones the pricing page shows. Yearly stickers are in the same catalog as 12 times a discounted monthly equivalent; confirm them on the pricing page toggle before you commit, because Stripe is the billing authority.

Sume plan stickers and concurrency, from the plan catalog and Generation admission docs, read 2026-10-03
PlanMonthlyYearly stickerConcurrent jobsQueue capacityAccepted jobs
Free$0none156
Pro$40$432 ($36/mo)42024
Startup$120$1,260 ($105/mo)84048
Scale$400$4,080 ($340/mo)20100120
EnterpriseCustomCustom20 default100 default120 default

What do concurrency and queue mean for my batch?

Concurrency is how many paid generation jobs can be processing at once. It is a dispatch limit, not a submit limit: a Free workspace can still submit six valid jobs and see five wait in queued while one runs. Queue capacity defaults to the larger of 3 and five times the concurrency, and accepted job capacity is the sum. Past that, a submit returns 429 queue_full and its reservation is released.

Prepaid top-ups do not raise concurrency; the docs say the limit is plan-only, with admin overrides for contract customers and a floor of 10 for organization workspaces. The effective number is always in the generation_limits block of a submit response, which is the figure to trust over any table.

How many API calls can I make per minute?

The API reference spells out the write budget per key by plan: Free 120 a minute, Pro 300, Startup 600 and Scale 1,200, with Enterprise by arrangement. The read budget defaults to 40 times the write budget, so a status-poll loop does not starve the submits that created the jobs. A 429 rate_limited carries details.scope to say which budget ran out, and retry-after tells you how long to wait.

Rate limits and concurrency are separate controls. A burst of 100 submits can pass the Pro write budget of 300 a minute and still land mostly in the queue, since only four run at a time.

  • Hitting concurrency: jobs wait in queued; no error.
  • Hitting queue capacity: 429 queue_full; wait or cancel queued jobs.
  • Hitting the write budget: 429 rate_limited; back off using retry-after.
  • Hitting the balance: 402 insufficient_credits; no job starts.

Which plan do I need for my workload?

Work backwards from job length. If a typical clip takes a few minutes to generate and you want 40 in an hour, the number that matters is how many can be in flight at once, which is the concurrency column, not the monthly price. A Pro workspace with four slots can drain a queue of 24 accepted jobs in six waves of at most four, and a Scale workspace with 20 slots does the same work in about two.

Spend is a separate question. Prices come from the rate card, for example $0.375 a second for a 768p H3 Max Recast or $0.01 a minute for transcription, and they are the same on every plan. Check the plan only for the three limits above and for access: the pricing page lists full Sume Agent access plus API, CLI and MCP access from Pro up, while Free offers image generation models and limited Agent access.

How do I read my own limits?

Do not hard-code the table. generation_limits comes back on each submit response with the effective concurrency_limit, queued_jobs_limit and queue_capacity_remaining, and GET /v1/balance returns the spendable USD. Together they tell you how much work to submit next. The docs also say wave_size_hint is only a submission hint and must not be used to size in-flight work, so size from concurrency_limit minus active and queued jobs instead.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume