Sume 429 queue_full from an MCP agent: wave sizes for Free to Scale

An agent that submits too many paid jobs hits 429 queue_full. Accepted capacity is 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Wave sizes inside.

5 min readSume
All posts

429 queue_full means the workspace has no accepted generation capacity left, so the agent submitted a wave larger than the plan allows. Accepted capacity is processing concurrency plus queue capacity: 6 jobs on Free, 24 on Pro, 48 on Startup and 120 on Scale and Enterprise. Wait for jobs to finish or cancel queued ones, then retry with the same idempotency_key.

Plan limits, read 2026-10-08
PlanProcessingQueueAccepted
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

A wave size to start with

The generation_admission_preview response carries wave_size_hint, which is max(1, floor(queue_capacity_remaining * 0.75)). The docs call it a submission hint only, not a concurrency limit. On an empty workspace with no admin override, remaining capacity equals accepted capacity, so the hint works out as below.

wave_size_hint on an empty workspace, computed from the plan table
PlanRemainingx 0.75Hint
Free64.54
Pro241818
Startup483636
Scale1209090

Agent loop

Teach the agent to ask before it fans out, then submit in waves.

  • Call generation_admission_preview and read queue_capacity_remaining and wave_size_hint.
  • Submit up to the hint, each with its own idempotency_key.
  • Call jobs_wait with job_ids for that wave.
  • Read results with a batch jobs_result.
  • Submit the next wave. On queue_full, wait, then retry with the same key.

Other errors with different fixes

429 rate_limited is request volume, not capacity: back off using retry-after. 402 insufficient_credits means the balance cannot reserve the estimate, and no provider work started. 409 idempotency_conflict means a key was reused for a different payload.

What changes the numbers

The table shows plan defaults. An admin override can change the effective limits, and the preview reports limit_source as plan or admin_override. When an override applies, do not size waves from the plan table; use concurrency_limit and queue_capacity_remaining from the response.

Retry behavior

On queue_full, do not hammer the endpoint. Wait for running jobs, or cancel queued ones you no longer need, then retry with the same key. A cancelled queued job frees capacity right away. Accepted work keeps its place, so a smaller wave submitted later is no slower overall than one large failing wave.

The arithmetic is simple, but only the preview knows the live state. Other jobs in the workspace, started by teammates or other agents, use the same capacity. If you see a hint of 4 on a Pro workspace, someone else has filled it. Ask the agent to report the numbers it read, so you can see whether your own wave caused the limit or a neighbor did.

Before you rely on this setup, run a short acceptance test with a read-only credential. Connect, call mcp_health, call tools_list, and read one job with jobs_status. Record the tool count you see, so you can notice later if a credential change alters it. Then repeat the test after any config edit. A five-minute test like this catches most wiring mistakes before they cost money, and it gives you a baseline to compare against when something behaves differently next week.

Sources

More in Integrations

All Integrations posts

Written by Sume