Cursor Projects coordinator fan-out: size waves from generation_limits

A Cursor coordinator that delegates to subagents can overrun a Sume workspace. Budget new in-flight jobs from generation_limits, not from wave_size_hint.

5 min readSume
All posts

A coordinator should size each wave of Sume jobs as max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. Cursor Projects (Sep 10, 2026) lets coordinator agents delegate to subagents, so the width of the fan-out is now a coordinator decision.

The Cursor changelog says coordinators share context across cloud and local, and can be triggered by Slack, schedules and PR streams.

The numbers to read

Generation submit responses include a generation_limits snapshot when Sume can compute one. These fields drive the budget.

generation_limits fields (read 2026-10-03)
FieldMeaning
concurrency_limitEffective maximum jobs in processing for the workspace.
active_generation_jobsJobs currently processing.
queued_generation_jobsJobs currently queued.
queue_capacity_remainingQueue budget plus idle processing seats before queue_full.
wave_size_hintSubmission-wave hint only; never a concurrency limit.

Worked examples from the docs

With concurrency_limit: 100, queued_jobs_limit: 500 and nothing running, queue_capacity_remaining is 600 and wave_size_hint is 450. The workspace is set to 100, so at most 100 new in-flight jobs fit in that snapshot. The 450 is only a hint that includes queue slots.

With 30 jobs processing and 10 queued, the new in-flight budget is 60. At zero headroom, wait and refresh the snapshot before submitting more.

A helper for the coordinator

This function turns a generation_limits object into a number of jobs the coordinator may start now. Count each newly submitted job against the budget until the next live snapshot.

def new_inflight_budget(limits: dict) -> int:
    concurrency = int(limits.get("concurrency_limit", 0))
    active = int(limits.get("active_generation_jobs", 0))
    queued = int(limits.get("queued_generation_jobs", 0))
    remaining = int(limits.get("queue_capacity_remaining", 0))
    return min(max(0, concurrency - active - queued), remaining)


limits = {
    "concurrency_limit": 100,
    "active_generation_jobs": 30,
    "queued_generation_jobs": 10,
    "queue_capacity_remaining": 560,
}
print(new_inflight_budget(limits))  # 60

Rules for the subagents

  • Each subagent submits with its own stable idempotency_key, so a retry returns the original job.
  • A queued job is normal. Subagents must not resubmit because a job is not yet processing.
  • On 429 queue_full, stop adding work, wait for jobs to finish, cancel queued jobs that are no longer needed, then retry with the same key.
  • Return job ids to the coordinator, which waits on batches of up to 20 with jobs_wait.

Triggers need caps

Because a schedule or a Slack message can start a coordinator without a person present, cap the work in the arguments. max_spend_usd is enforced only when provided, and dry_run=true previews cost without a job. Use generation_admission_preview before an expensive burst.

If a reviewer only reads outputs, give it a read-only OAuth session. Paid tools return insufficient_scope without mcp:write.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume