How many Wan 3.0 clips run at once on Sume: 1, 4, 8 or 20
Wan 3.0 jobs run under plan-only concurrency: Free 1, Pro 4, Startup 8, Scale 20. Extra jobs queue. What that means for a 30-second render batch.

On Sume, wan-3.0 jobs run under the same plan-only processing cap as every paid generation: 1 at once on Free, 4 on Pro, 8 on Startup, and 20 on Scale and Enterprise. Valid jobs beyond the cap are accepted as queued until a slot opens. Buying more prepaid balance does not raise the cap.
Limits come from Generation admission, read 2026-10-02; Wan 3.0's 2-30 s range is from the Video Router docs. TechNode reports that Wan 3.0 launched on 2026-08-24 with 30-second generation (search summary only, read 2026-10-02).
What does the cap mean for a batch of Wan clips?
Wall-clock time depends on how long each job runs, which Sume does not publish as a fixed number, so plan in waves: the number of waves is your clip count divided by the processing cap, rounded up. Sume exposes queue counts and remaining capacity, not a per-job ETA.
A batch of 40 clips on Pro runs in 10 waves; on Scale it runs in 2.
| Plan | Processing at once | Waves for 40 clips | Accepted at once |
|---|---|---|---|
| Free | 1 | 40 | 6 |
| Pro | 4 | 10 | 24 |
| Startup | 8 | 5 | 48 |
| Scale | 20 | 2 | 120 |
Can I submit all 40 at once?
Only up to accepted capacity: 24 on Pro, 48 on Startup. Beyond that a submit fails with 429 queue_full. Use the generation_limits snapshot in each submit response and stop adding work when queue_capacity_remaining is low.
The docs warn against treating wave_size_hint as concurrency; it is a submission hint, and the effective cap is concurrency_limit. The dashboard Concurrency tab is the source of truth for a workspace.
How do I keep a long Wan batch safe to retry?
Send a distinct Idempotency-Key per clip and reuse it only for exact retries, and submit with mode: "async". Sync mode waits at most 30 seconds, which a long Wan clip may not fit in.
Store each job id, poll with backoff, and do not resubmit just because your worker timed out. For cost per tier see Wan 3.0 cost for 100 clips.
Before a large run, read the generation_limits block from your first submit response and compute headroom as the docs describe. If headroom is zero, wait. Pacing submissions to your headroom keeps you under the limit, so you avoid queue_full responses.
Sources
Related posts
More in Developers
- Wan 3.0 job failed: read the error category, then pick the next action
When a wan-3.0 job fails, Sume exposes a category, retryability and next action. Which categories to fix, retry or escalate, with the 409 and 429 codes.
- Wan 3.0 request returns 409 idempotency_conflict after a prompt edit
Reusing an Idempotency-Key with a changed Wan 3.0 payload returns 409 idempotency_conflict. Use a new key for a new prompt; reuse it only for exact retries.
- Wan 3.0 polling every 15 seconds vs Sume's 30-second poll
Alibaba recommends polling Wan 3.0 every 15 seconds. Sume's video guide uses 30 seconds and a signed callback_url. Which to pick and why.
- Wan 3.0 has six regional endpoints; Sume has one base URL
Alibaba's Wan 3.0 API is served from six regions with a workspace id in the URL. Sume's video API uses api.sume.com/v1/videos with no region field.
Written by Sume