Claude Code 2.1.289 agent.spawn: teammates share one Sume queue

Claude Code 2.1.289 adds agent.spawn for teammates. Teammates sharing a Sume workspace share its concurrency limit and queue, so plan the width of the fan-out.

4 min readSume
All posts

Teammates that call Sume from the same workspace share one generation concurrency limit and one queue, however many agents you spawn. The Claude Code changelog lists agent.spawn for teammates in 2.1.289 (Oct 3, 2026), which makes wide fan-out easy and makes the shared limit worth planning for.

Sume's limits are per workspace, not per agent, so spawning more teammates does not buy more parallel renders.

The limits a team shares

The docs say to prefer the effective generation_limits.concurrency_limit field over this static table, because admin overrides can raise it. Concurrency is plan-only; prepaid top-ups do not raise it.

Default generation limits by plan, from the admission docs (read 2026-10-03)
PlanProcessing concurrencyQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

What happens when teammates overshoot

Concurrency is a dispatch limit, not a submit limit. Valid jobs beyond the processing cap are accepted as queued and move to processing when a slot opens. Only when queue capacity is also consumed does a submit fail, with 429 queue_full.

That means a team of teammates will mostly see jobs sit in queued, which is normal, not a failure. A teammate that treats queued as an error and resubmits makes the queue fuller and the bill larger.

Give the lead agent the budget

A simple pattern is to let one lead agent own the sizing and hand each teammate a slice of work.

  • Read generation_limits from a submit response, or use generation_admission_preview first.
  • Budget new in-flight work as max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining.
  • Do not use wave_size_hint as a concurrency number; the docs call it a submission-wave hint only.
  • Have every teammate use its own stable idempotency_key per paid call, so a retry returns the original job.

Collect results in one wait

When several teammates each own a job, the lead can wait on all the ids together. jobs_wait takes 1 to 20 job_ids and wait_for of all or any, and include_results: true returns finished results in the same answer.

Each call holds at most 55 seconds. If it returns wait_slice_expired, repeat it with the same ids rather than submitting again.

Scope the credential too

Under OAuth, mcp:read sessions see read-only tools, and paid tools return insufficient_scope until mcp:write is granted. A reviewer teammate that only inspects job results does not need write access. Giving it a read-only session keeps a spawned agent from spending.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume