Free plan seventh submit: 429 queue_full is free, 402 can come first
On Sume Free, six paid jobs fit (1 processing, 5 queued). The seventh gets 429 queue_full with its hold released. Which error comes first, and the cost.

On the Free plan the seventh paid generation submit is rejected with 429 queue_full, and the rejection costs nothing: the admission docs say Sume releases or refunds the reservation for a failed admission. Free processes one job at a time and queues five more, so six accepted jobs is the ceiling. A 402 insufficient_credits is a different failure, and it can arrive first if the balance cannot cover the hold.
Both errors happen before provider work starts, so neither bills a generation. What differs is the fix: queue_full is solved by waiting, 402 by adding funds or sending a cheaper request.
Accepted capacity by plan
Concurrency is plan-only; prepaid top-ups do not raise it. The default queue is max(3, concurrency x 5), and accepted capacity is concurrency plus queue.
| Plan | Processing | Queue | Accepted jobs | Seventh submit |
|---|---|---|---|---|
| Free | 1 | 5 | 6 | 429 queue_full |
| Pro | 4 | 20 | 24 | accepted |
| Startup | 8 | 40 | 48 | accepted |
| Scale | 20 | 100 | 120 | accepted |
Worked example with 5-second Omni clips
Seven gemini-omni-flash-1.1 clips of 5 seconds at 720p each cost 5 x $0.125 = $0.625. Six accepted jobs hold $3.75. With a $5.00 balance the seventh submit (it would need $4.375 in total) fails on admission for capacity, and the hold for that attempt is released. With a $3.00 balance the fifth submit already fails with a 402 because the balance cannot reserve a fifth $0.625 on top of $2.50 held.
So the order in which a client sees the two errors depends on the balance and the plan. A script should read the code, not the status family: both are 4xx, and a retry loop that treats them alike will hammer a workspace that needs money rather than time.
What to do on each
For queue_full, stop adding work, poll the running jobs until one reaches a terminal state, cancel queued jobs you no longer need, and retry with the same idempotency key once capacity opens. Use retry-after if it is present. For 402, follow the docs: upgrade or wait for included credit, or send a smaller request.
429 rate_limited is a third case. It means request volume, not capacity or money, and it is the one where backoff on retry-after is the documented answer. Read-only status polling has its own limit and is not generation concurrency.
Writing the client
Branch on error.code, not on the status number. On queue_full, sleep until a polled job reaches a terminal state, then resubmit with the same key. On insufficient_credits, stop and alert, because waiting will not help. On rate_limited, back off by retry-after. Keep the request id from the error body in your logs; it is safe to share with support. If you submit through MCP, jobs_wait holds at most 55 seconds per call, so loop on wait_slice_expired and never resubmit the paid create.
Sources
Related posts
More in Pricing
- Free plan accepts six jobs: six 30-second Wan 3.0 clips hold $11.25
Free workspaces run 1 job and queue 5. Six 30-second Wan 3.0 480p jobs reserve $11.25; the seventh gets 429 queue_full. Pro, Startup and Scale caps shown.
- A full queue on each Sume plan: reserved dollars for 10-second jobs
If every accepted job holds its estimate, a full Scale queue of 120 ten-second Seedance 2.5 jobs at 1080p reserves $1,705.86. Free, Pro and Startup too.
- Gemini 3.8 Flash-Lite TTS: 30,000 ten-second product clips cost $45
30,000 ten-second product voice clips are 250 audio tokens each: $45 on Flash-Lite and $67.50 on Flash in 2026, $90 and $135 from January 1, before text input.
- Gemini 3.8 TTS for a 17-minute weekly audio newsletter: 2026 vs 2027
A 17-minute newsletter is 25,500 audio tokens: $0.23 per issue on Gemini 3.8 Flash TTS in 2026, $0.46 from January 1. Year totals for Flash and Flash-Lite.
Written by Sume