Free plan accepts six jobs: six 30-second Wan 3.0 clips hold $11.25
Free workspaces run 1 job and queue 5. Six 30-second Wan 3.0 480p jobs reserve $11.25; the seventh gets 429 queue_full. Pro, Startup and Scale caps shown.

A Free workspace on Sume can have six paid generation jobs open at once: one processing and five queued. A seventh submit fails with 429 queue_full. If all six are 30-second Wan 3.0 clips at 480p, together they hold 6 x $1.875 = $11.25 of balance.
The caps by plan
Sume separates processing concurrency from queue capacity. Concurrency is plan-only; prepaid top-ups do not raise it. Queue capacity defaults to max(3, concurrency x 5), and accepted capacity is concurrency plus queue. The table below applies each plan's accepted capacity to a 30-second Wan 3.0 480p job at $1.875.
| Plan | Processing | Queue | Accepted | Reserve if all full |
|---|---|---|---|---|
| Free | 1 | 5 | 6 | $11.25 |
| Pro | 4 | 20 | 24 | $45.00 |
| Startup | 8 | 40 | 48 | $90.00 |
| Scale | 20 | 100 | 120 | $225.00 |
What it changes in practice
Full concurrency is not an error. On Free you can submit six jobs together and they will run one at a time; the Free plan finishes six 30-second jobs in sequence, not in parallel. The docs advise treating queued as normal and polling with backoff.
The cap also limits how much money can be locked. Your own balance may be the tighter limit: at $11.25 per six jobs, a $10 balance accepts only five of them (5 x 1.875 = $9.375) and the sixth is refused with 402 insufficient_credits rather than queue_full.
- 429 queue_full: capacity is full; wait or cancel queued jobs, then retry with the same Idempotency-Key.
- 402 insufficient_credits: the balance cannot cover the reserve.
- The dashboard Concurrency tab is the source of truth; org workspaces have a floor of 10.
Time, not only money
With one processing slot, six jobs run in sequence. If a 30-second Wan 3.0 job takes a few minutes to generate, the sixth job waits for five others. Sume does not publish a per-job ETA or queue position, so plan on polling, not on a promised finish time. Pro runs four jobs at once, so six jobs there finish in roughly two waves instead of six.
If speed matters more than the plan price, compare the cost of the plan with what the wait costs you. The generation docs say that concurrency is plan-only and that prepaid top-ups do not raise it, so adding balance will not make a Free workspace faster.
Sizing a batch
Use the generation_limits object in submit responses. The docs define the budget for new in-flight work as max(0, concurrency_limit - active - queued), limited to queue_capacity_remaining. Details: generation admission and errors. The plan figures are the docs' defaults and the effective values can differ.
Sources
Related posts
More in Pricing
- A full queue on each Sume plan: reserved dollars for 10-second jobs
If every accepted job holds its estimate, a full Scale queue of 120 ten-second Seedance 2.5 jobs at 1080p reserves $1,705.86. Free, Pro and Startup too.
- Gemini 3.8 Flash-Lite TTS: 30,000 ten-second product clips cost $45
30,000 ten-second product voice clips are 250 audio tokens each: $45 on Flash-Lite and $67.50 on Flash in 2026, $90 and $135 from January 1, before text input.
- Gemini 3.8 TTS for a 17-minute weekly audio newsletter: 2026 vs 2027
A 17-minute newsletter is 25,500 audio tokens: $0.23 per issue on Gemini 3.8 Flash TTS in 2026, $0.46 from January 1. Year totals for Flash and Flash-Lite.
- Gemini 3.8 TTS, 500 12-second reads: $2.43 priority, $0.68 batch
500 reads of 12 seconds is 150,000 audio tokens. Gemini 3.8 Flash TTS costs $1.35 standard, $2.43 priority, $0.675 batch or flex in 2026. Sume TTS is $4.28.
Written by Sume