How many Wan 3.0 clips run at once on Sume: 1, 4, 8 or 20

Wan 3.0 jobs run under plan-only concurrency: Free 1, Pro 4, Startup 8, Scale 20. Extra jobs queue. What that means for a 30-second render batch.

4 min readSume
All posts

On Sume, wan-3.0 jobs run under the same plan-only processing cap as every paid generation: 1 at once on Free, 4 on Pro, 8 on Startup, and 20 on Scale and Enterprise. Valid jobs beyond the cap are accepted as queued until a slot opens. Buying more prepaid balance does not raise the cap.

Limits come from Generation admission, read 2026-10-02; Wan 3.0's 2-30 s range is from the Video Router docs. TechNode reports that Wan 3.0 launched on 2026-08-24 with 30-second generation (search summary only, read 2026-10-02).

What does the cap mean for a batch of Wan clips?

Wall-clock time depends on how long each job runs, which Sume does not publish as a fixed number, so plan in waves: the number of waves is your clip count divided by the processing cap, rounded up. Sume exposes queue counts and remaining capacity, not a per-job ETA.

A batch of 40 clips on Pro runs in 10 waves; on Scale it runs in 2.

Waves for a 40-clip batch from the plan caps in Sume's Generation admission docs, read 2026-10-02.
PlanProcessing at onceWaves for 40 clipsAccepted at once
Free1406
Pro41024
Startup8548
Scale202120

Can I submit all 40 at once?

Only up to accepted capacity: 24 on Pro, 48 on Startup. Beyond that a submit fails with 429 queue_full. Use the generation_limits snapshot in each submit response and stop adding work when queue_capacity_remaining is low.

The docs warn against treating wave_size_hint as concurrency; it is a submission hint, and the effective cap is concurrency_limit. The dashboard Concurrency tab is the source of truth for a workspace.

How do I keep a long Wan batch safe to retry?

Send a distinct Idempotency-Key per clip and reuse it only for exact retries, and submit with mode: "async". Sync mode waits at most 30 seconds, which a long Wan clip may not fit in.

Store each job id, poll with backoff, and do not resubmit just because your worker timed out. For cost per tier see Wan 3.0 cost for 100 clips.

Before a large run, read the generation_limits block from your first submit response and compute headroom as the docs describe. If headroom is zero, wait. Pacing submissions to your headroom keeps you under the limit, so you avoid queue_full responses.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume