Free plan vs paid wallet for AI video: what each one buys on Sume

On Sume the plan sets how many video jobs run at once and your request rates; the wallet pays for each clip. A top-up never raises concurrency.

5 min readSume
All posts

On Sume, the plan and the wallet buy different things. The plan sets how many video jobs run at the same time and how many API requests you can send per minute. The wallet pays for each clip. A top-up does not increase processing concurrency, so a bigger balance on the Free plan still renders one job at a time.

What the plan controls

The admission docs say generation concurrency is plan-only: "Prepaid top-ups do not increase the processing concurrency limit." The same plan also sets the API request budget in the authentication docs.

What the plan sets (Sume docs, read 2026-10-05)
PlanConcurrent jobsQueue capacityAccepted jobsWrites per minuteReads per minute
Free1561204800
Pro4202430012000
Startup8404860024000
Scale20100120120048000

What the wallet controls

The wallet is one USD balance. Video generation, Sume Agent, Formats and the API all draw from it at each model's published rate. The pricing FAQ describes on-demand purchases from $10 to $1000 for a shared wallet. A submit reserves the estimated cost first; if the balance cannot cover it, the call fails with 402 insufficient_credits before provider work starts.

Which one limits you first

Most people hit the plan limit before the wallet limit when they batch. With Free concurrency of 1, a Wan 3.0 clip at 480p for 5 seconds bills $0.32 and a $10 balance covers 31 of them, but they render one after another. On Pro, four render together. The queue still accepts extra jobs while capacity remains: Free accepts 6 jobs in total, then 429 queue_full.

So the rule of thumb is: size the plan for how many clips you need in flight, and size the wallet for how many clips you need in total.

Do not mix up the errors

402 is the wallet: add funds or send a cheaper request. 429 queue_full is the plan: wait for jobs to finish or cancel queued ones. 429 rate_limited is request volume, and it carries retry-after. Treat each one differently, because retrying a 402 never helps until the balance changes.

Reading the live numbers

Do not hard-code the table. The docs say to prefer the effective field to the static table: generation_limits.concurrency_limit in a submit response, and the **Concurrency** tab in the dashboard. Admin overrides can raise the effective number, and org workspaces have a floor of 10. The response also tells you queue_capacity_remaining, which is the remaining queued-job budget plus idle processing seats before queue_full.

A simple preflight is two reads: GET /v1/balance for the wallet and the last generation_limits for the plan. If the balance covers the batch and queue_capacity_remaining is at least the batch size, submit. If only one holds, fix that one first.

Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.

A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume