Claude Code 2.1.289 agent.spawn: teammates share one Sume queue
Claude Code 2.1.289 adds agent.spawn for teammates. Teammates sharing a Sume workspace share its concurrency limit and queue, so plan the width of the fan-out.

Teammates that call Sume from the same workspace share one generation concurrency limit and one queue, however many agents you spawn. The Claude Code changelog lists agent.spawn for teammates in 2.1.289 (Oct 3, 2026), which makes wide fan-out easy and makes the shared limit worth planning for.
Sume's limits are per workspace, not per agent, so spawning more teammates does not buy more parallel renders.
The limits a team shares
The docs say to prefer the effective generation_limits.concurrency_limit field over this static table, because admin overrides can raise it. Concurrency is plan-only; prepaid top-ups do not raise it.
| Plan | Processing concurrency | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
What happens when teammates overshoot
Concurrency is a dispatch limit, not a submit limit. Valid jobs beyond the processing cap are accepted as queued and move to processing when a slot opens. Only when queue capacity is also consumed does a submit fail, with 429 queue_full.
That means a team of teammates will mostly see jobs sit in queued, which is normal, not a failure. A teammate that treats queued as an error and resubmits makes the queue fuller and the bill larger.
Give the lead agent the budget
A simple pattern is to let one lead agent own the sizing and hand each teammate a slice of work.
- Read
generation_limitsfrom a submit response, or usegeneration_admission_previewfirst. - Budget new in-flight work as
max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped byqueue_capacity_remaining. - Do not use
wave_size_hintas a concurrency number; the docs call it a submission-wave hint only. - Have every teammate use its own stable
idempotency_keyper paid call, so a retry returns the original job.
Collect results in one wait
When several teammates each own a job, the lead can wait on all the ids together. jobs_wait takes 1 to 20 job_ids and wait_for of all or any, and include_results: true returns finished results in the same answer.
Each call holds at most 55 seconds. If it returns wait_slice_expired, repeat it with the same ids rather than submitting again.
Scope the credential too
Under OAuth, mcp:read sessions see read-only tools, and paid tools return insufficient_scope until mcp:write is granted. A reviewer teammate that only inspects job results does not need write access. Giving it a read-only session keeps a spawned agent from spending.
Sources
Related posts
More in Agents
- Cursor Projects coordinator fan-out: size waves from generation_limits
A Cursor coordinator that delegates to subagents can overrun a Sume workspace. Budget new in-flight jobs from generation_limits, not from wave_size_hint.
- Parallel tool calls on a voice agent: one idempotency key each
ElevenLabs agents default enable_parallel_tool_calls to true. If tools start paid jobs, give each call its own key and wait on all job ids in one batch.
- OpenAI Dots x Runway: brief a dot, it plans shots. The Sume loop
Runway's Oct 1 changelog lists OpenAI Dots x Runway for all plans: brief a dot, it plans shots, Runway shoots. The same loop with Sume MCP tools.
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
Written by Sume