Sume queue_full: is money held for the job that was rejected?
A 429 queue_full means the workspace has no accepted-job capacity left. Sume releases or refunds the reservation for the failed admission. How to confirm it.

Short answer
No money should stay held for a job that returned 429 queue_full. The generation admission docs say that, when applicable, Sume releases or refunds the reservation for the failed admission (Generation admission). The usage summary also has a refunded_usd_micros field described as holds Sume gave back after a failure, cancellation or queue_full. It is not spend.
When queue_full happens
Concurrency is a dispatch limit, not a submit limit. Sume accepts valid jobs as queued while queue capacity remains. queue_full appears only when the accepted capacity is used up: concurrency plus queue. By default the queue holds max(3, concurrency_limit x 5) jobs.
For the plans in the docs that works out as follows.
| Plan | Processing | Queue | Accepted |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
Confirm the money came back
The error body can include a generation_limits snapshot and job metadata for the failed attempt. To check the wallet, read the balance before and after the burst and compare with the ledger.
Row status refunded means the reservation was released. If a rejected submit left held_usd_micros open for more than a moment, send the request id from the error envelope to Sume support. The envelope carries a request_id that is safe to share.
curl https://api.sume.com/v1/balance \
-H "Authorization: Bearer $SUME_API_KEY"
curl "https://api.sume.com/v1/usage?limit=20" \
-H "Authorization: Bearer $SUME_API_KEY"What to do on queue_full
- Stop adding work for that workspace.
- Poll current jobs until one reaches a terminal state, and cancel queued jobs you do not need.
- Retry with the same
Idempotency-Keyafter capacity opens; useretry-afterwhen it is present. - Do not confuse it with
429 rate_limited, which is request volume rather than generation capacity.
A worked example on the Free plan
On the Free plan the accepted capacity is 6: one processing and five queued. Submit eight 10-second Wan 720p jobs at once. The first six are accepted and hold $1.25 each, $7.50 in total. The seventh and eighth return 429 queue_full. Neither should add to the held amount, and the docs say the reservation for a failed admission is released or refunded.
When the first job completes, a queue slot opens. Retrying the seventh with the same idempotency key then succeeds, and the held amount rises by $1.25 again.
Common mistakes
Two mistakes recur. The first is treating queue_full like a rate limit and retrying in a tight loop; the docs tell you to wait for a job to finish or to cancel queued work. The second is using wave_size_hint as a concurrency limit. It is a submission-wave hint only. Size in-flight work from concurrency_limit and the live counts in generation_limits.
Pay attention to the snapshot nature of these counts. They can change right after the response when workers claim jobs or other clients submit work.
The check that matters is the ledger, not the response alone. After a burst of submits, read the usage summary and confirm the held amount equals the sum of the jobs you know were accepted. If it is higher, look for jobs you did not mean to submit, such as a worker that retried without an idempotency key. If it matches, the rejected submits cost nothing and you can retry them as capacity opens.
A steady approach for large batches is to submit up to the accepted capacity, poll, and top the queue up as jobs finish. That keeps the wallet hold close to the work actually in flight and avoids a wall of 429 responses.
Sources
Related posts
More in Pricing
- Quiz audio: 100 questions at 140 characters, one job, sliced per line
A hundred 140-character quiz questions are 14,000 characters: $0.21 on MAI Flash, $0.31 on MAI-Voice-2.1 and $0.67 on Sume, which can slice one job by sentence.
- Quote a 30-second AI video clip: Seedance 2.5 per clip and per minute
How to quote a 30-second Seedance 2.5 clip on Sume: price at 480p, 720p and 1080p, a per-minute equivalent, a retake allowance and the rounding rule.
- Re-render your top old videos first under a fixed $3 budget
Rank clips from a retired video pipeline by views per dollar of Sume re-render cost, then fill a fixed budget greedily. Python, arithmetic shown.
- Re-render a gpt-image-1 library before Oct 23: cost on 2.5
gpt-image-1 shuts down Oct 23, 2026. Re-rendering 1,000 images on GPT Image 2.5 at 1024x1024 costs $7.35 at low, $16.46 at medium and $65.85 at high on Sume.
Written by Sume