Sume queue_full: is money held for the job that was rejected?

A 429 queue_full means the workspace has no accepted-job capacity left. Sume releases or refunds the reservation for the failed admission. How to confirm it.

4 min readSume
All posts

Short answer

No money should stay held for a job that returned 429 queue_full. The generation admission docs say that, when applicable, Sume releases or refunds the reservation for the failed admission (Generation admission). The usage summary also has a refunded_usd_micros field described as holds Sume gave back after a failure, cancellation or queue_full. It is not spend.

When queue_full happens

Concurrency is a dispatch limit, not a submit limit. Sume accepts valid jobs as queued while queue capacity remains. queue_full appears only when the accepted capacity is used up: concurrency plus queue. By default the queue holds max(3, concurrency_limit x 5) jobs.

For the plans in the docs that works out as follows.

Accepted job capacity by plan (Sume docs, read 2026-10-05)
PlanProcessingQueueAccepted
Free156
Pro42024
Startup84048
Scale20100120

Confirm the money came back

The error body can include a generation_limits snapshot and job metadata for the failed attempt. To check the wallet, read the balance before and after the burst and compare with the ledger.

Row status refunded means the reservation was released. If a rejected submit left held_usd_micros open for more than a moment, send the request id from the error envelope to Sume support. The envelope carries a request_id that is safe to share.

curl https://api.sume.com/v1/balance \
  -H "Authorization: Bearer $SUME_API_KEY"

curl "https://api.sume.com/v1/usage?limit=20" \
  -H "Authorization: Bearer $SUME_API_KEY"

What to do on queue_full

  • Stop adding work for that workspace.
  • Poll current jobs until one reaches a terminal state, and cancel queued jobs you do not need.
  • Retry with the same Idempotency-Key after capacity opens; use retry-after when it is present.
  • Do not confuse it with 429 rate_limited, which is request volume rather than generation capacity.

A worked example on the Free plan

On the Free plan the accepted capacity is 6: one processing and five queued. Submit eight 10-second Wan 720p jobs at once. The first six are accepted and hold $1.25 each, $7.50 in total. The seventh and eighth return 429 queue_full. Neither should add to the held amount, and the docs say the reservation for a failed admission is released or refunded.

When the first job completes, a queue slot opens. Retrying the seventh with the same idempotency key then succeeds, and the held amount rises by $1.25 again.

Common mistakes

Two mistakes recur. The first is treating queue_full like a rate limit and retrying in a tight loop; the docs tell you to wait for a job to finish or to cancel queued work. The second is using wave_size_hint as a concurrency limit. It is a submission-wave hint only. Size in-flight work from concurrency_limit and the live counts in generation_limits.

Pay attention to the snapshot nature of these counts. They can change right after the response when workers claim jobs or other clients submit work.

The check that matters is the ledger, not the response alone. After a burst of submits, read the usage summary and confirm the held amount equals the sum of the jobs you know were accepted. If it is higher, look for jobs you did not mean to submit, such as a worker that retried without an idempotency key. If it matches, the rejected submits cost nothing and you can retry them as capacity opens.

A steady approach for large batches is to submit up to the accepted capacity, poll, and top the queue up as jobs finish. That keeps the wallet hold close to the work actually in flight and avoids a wall of 429 responses.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume