Sume plans: Pro $40, Startup $120, Scale $400 and what they limit
What each Sume plan sets: monthly price, concurrent jobs, queue size and API write budget. Usage is billed at each model's rate, not by the plan.

Sume's self-serve plans are Pro at $40 a month, Startup at $120 and Scale at $400, with a Free plan at $0 and Enterprise by contact. What a plan changes is access and limits: how many generation jobs run at once, how many can wait in the queue, and how many API writes you can make per minute. Generation itself is billed per use at each model's published rate from your USD balance, so the plan price is not a rate card.
If you are comparing against credit-bucket tools, that is the main difference to hold on to: Sume's pricing page describes the plan grid as product access and capabilities and says usage is billed at each model's published rate. The tables below list the limits from Sume's docs and plan catalog, and then show how to work out which plan you need from your own job pattern.
What does each plan cost and include?
Monthly prices come from the plan catalog that the public pricing page and the dashboard both read. The feature lines are the ones the pricing page shows. Yearly stickers are in the same catalog as 12 times a discounted monthly equivalent; confirm them on the pricing page toggle before you commit, because Stripe is the billing authority.
| Plan | Monthly | Yearly sticker | Concurrent jobs | Queue capacity | Accepted jobs |
|---|---|---|---|---|---|
| Free | $0 | none | 1 | 5 | 6 |
| Pro | $40 | $432 ($36/mo) | 4 | 20 | 24 |
| Startup | $120 | $1,260 ($105/mo) | 8 | 40 | 48 |
| Scale | $400 | $4,080 ($340/mo) | 20 | 100 | 120 |
| Enterprise | Custom | Custom | 20 default | 100 default | 120 default |
What do concurrency and queue mean for my batch?
Concurrency is how many paid generation jobs can be processing at once. It is a dispatch limit, not a submit limit: a Free workspace can still submit six valid jobs and see five wait in queued while one runs. Queue capacity defaults to the larger of 3 and five times the concurrency, and accepted job capacity is the sum. Past that, a submit returns 429 queue_full and its reservation is released.
Prepaid top-ups do not raise concurrency; the docs say the limit is plan-only, with admin overrides for contract customers and a floor of 10 for organization workspaces. The effective number is always in the generation_limits block of a submit response, which is the figure to trust over any table.
How many API calls can I make per minute?
The API reference spells out the write budget per key by plan: Free 120 a minute, Pro 300, Startup 600 and Scale 1,200, with Enterprise by arrangement. The read budget defaults to 40 times the write budget, so a status-poll loop does not starve the submits that created the jobs. A 429 rate_limited carries details.scope to say which budget ran out, and retry-after tells you how long to wait.
Rate limits and concurrency are separate controls. A burst of 100 submits can pass the Pro write budget of 300 a minute and still land mostly in the queue, since only four run at a time.
- Hitting concurrency: jobs wait in
queued; no error. - Hitting queue capacity:
429 queue_full; wait or cancel queued jobs. - Hitting the write budget:
429 rate_limited; back off usingretry-after. - Hitting the balance:
402 insufficient_credits; no job starts.
Which plan do I need for my workload?
Work backwards from job length. If a typical clip takes a few minutes to generate and you want 40 in an hour, the number that matters is how many can be in flight at once, which is the concurrency column, not the monthly price. A Pro workspace with four slots can drain a queue of 24 accepted jobs in six waves of at most four, and a Scale workspace with 20 slots does the same work in about two.
Spend is a separate question. Prices come from the rate card, for example $0.375 a second for a 768p H3 Max Recast or $0.01 a minute for transcription, and they are the same on every plan. Check the plan only for the three limits above and for access: the pricing page lists full Sume Agent access plus API, CLI and MCP access from Pro up, while Free offers image generation models and limited Agent access.
How do I read my own limits?
Do not hard-code the table. generation_limits comes back on each submit response with the effective concurrency_limit, queued_jobs_limit and queue_capacity_remaining, and GET /v1/balance returns the spendable USD. Together they tell you how much work to submit next. The docs also say wave_size_hint is only a submission hint and must not be used to size in-flight work, so size from concurrency_limit minus active and queued jobs instead.
Sources
Related posts
More in Pricing
- TTS cost by character: the same sentence in four languages
Sume bills TTS per character, not per second. One sentence counted in English, French, German and Korean, with the 1-cent floor and the 20,000-character cap.
- Which Sume API calls are free: balance, usage, catalog, filter check
Balance, usage and catalog reads cost nothing on Sume, and neither does the video-filter check. What is billed, what is refunded, and what a 402 means.
- How Sume pricing works: plans, one wallet, published model rates
Sume plans set access and concurrency. Usage draws from one prepaid wallet at each model's published USD rate, for generation, the Agent, Formats, and the API.
- AI avatar video API pricing: cost per second and per minute
Sume bills AI avatar video per second by quality tier, with separate rates when you send a product image. Per-minute costs for standard, plus, and max.
Written by Sume