Eight Wan 3.0 clips on Free: six queue, two get queue_full
Free plan accepted capacity is 6 paid jobs. Submit eight 10-second Wan 3.0 720p jobs and two return 429 queue_full; the six accepted hold $7.50 of the $10.00.

On the Free plan, six of eight simultaneous Wan 3.0 jobs are accepted and two come back as 429 queue_full. The Free plan processes 1 job at a time and queues 5 more, so accepted capacity is 1 + 5 = 6. Eight 10-second Wan 3.0 jobs at 720p are 10 x $0.125 = $1.25 each. The six accepted jobs hold $7.50; the two rejected ones hold nothing, because Sume refunds holds after queue_full.
Free plan limits against the batch
The figures come from the generation admission page. Concurrency is a dispatch limit, not a submit limit: a valid job is accepted as queued while queue capacity remains. The default queue capacity is max(3, concurrency x 5).
| Item | Value | Arithmetic |
|---|---|---|
| Processing concurrency | 1 | Free plan |
| Queue capacity | 5 | max(3, 1 x 5) |
| Accepted job capacity | 6 | 1 + 5 |
| Jobs submitted | 8 | your batch |
| Accepted as queued or processing | 6 | min(8, 6) |
| Rejected with 429 queue_full | 2 | 8 - 6 |
| Price per job | $1.25 | 10 s x $0.125 |
| Open holds | $7.50 | 6 x $1.25 |
What to do with the two rejected jobs
queue_full is not a request-rate error. Sume's errors page says it means no new paid generation job can be accepted until a queued or processing job finishes or is canceled. Wait for a slot, then resubmit the same two requests. Use an Idempotency-Key on every submit so that a retry whose first response you never saw returns the original job and does not create a second paid job.
Do not treat queued as a failure. Store the job id and poll with backoff, or send a webhook URL and wait for the terminal event.
A retry loop that respects the limit
A simple client keeps a list of pending requests, submits until it sees queue_full, stops, and waits for the oldest job to reach a terminal state before sending the next. Because every request carries its own Idempotency-Key, a submit whose response was lost can be sent again safely. The rate-limit headers ratelimit-remaining and retry-after are separate signals for request volume, not for queue space.
On the Free plan the wall-clock effect is also worth planning for. Only one job processes at a time, so the sixth accepted job waits behind five others. If you need all eight clips in parallel, the upgrade to Pro, which processes 4 at once and accepts 24, is the lever; prepaid balance does not change concurrency.
The cheaper fix
Waves are the standard answer: submit up to the accepted capacity, wait for jobs to finish, submit the next wave. For eight clips on Free that is a wave of 6 and a wave of 2. A higher plan removes the limit: Pro accepts 24 at once, Startup 48 and Scale 120 per the same page. The dashboard Concurrency tab shows your workspace's effective limit, and the page says to prefer that field to the static table.
Sources
Related posts
More in Developers
- A 7-second clip on six Sume video models: $0.525 to $4.0446
Per-second list prices times 7 s for six Sume video models, with each model's accepted duration range and a Python check that rejects out-of-range lengths.
- A fresh UUID per retry is not an idempotency key (Node, Sume)
Create the Idempotency-Key once outside the retry loop, retry only 429, 502, 503 and 504, and stop on 4xx. A Node submit helper for Sume /v1/videos.
- Per-request timeout on Sume video polls: 15 s abort, then poll again
Give every poll GET its own 15 s AbortSignal.timeout so one hung read does not freeze a 30 s Wan or Seedance job loop. Reads per job per hour at 5, 10, 30 s.
- Agent Completion webhook retries: 10 attempts over about 3 hours
Sume tries an agent.run.terminal webhook up to 10 times. By the documented formula the nine waits add up to about 3 h 3 min, before jitter and Retry-After.
Written by Sume