Three Responses subagents, one Sume workspace queue
Responses multi_agent runs 3 subagents by default and they share your tools. Sume accepts extra jobs as queued until queue_full; give each create its own key.

Three subagents can submit video jobs at once, and Sume will accept them: concurrency is a dispatch limit, not a submit limit. Jobs beyond your processing concurrency wait as queued until the queue is full, then new paid submissions fail with 429 queue_full.
The OpenAI multi-agent guide, read 2026-09-30, says max_concurrent_subagents sets the maximum number of subagents active at once across the whole agent tree, defaults to 3, and that all subagents share the request's model and available tools. If a Sume tool is in that list, every subagent can call it.
How many jobs can my workspace hold?
Per Generation admission, processing concurrency is plan-based, and queue capacity defaults to the larger of 3 and five times the concurrency limit. The dashboard Concurrency tab and the effective generation_limits.concurrency_limit field are the source of truth; the table is the static default.
| Plan | Processing | Queue | Accepted |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
What should each subagent send?
Each paid create needs its own idempotency_key; the field is required on write and paid tools. Derive it from the subagent's task, not from the request, so a retry of the same task reuses the key and a different task never collides. A Free workspace accepts 6 jobs in total, so three subagents each starting two videos fills it.
How do I collect the results?
Prefer one batch wait after the fan-out instead of N single waits: jobs_wait takes 1 to 20 job_ids and an optional wait_for of all or any. On wait_slice_expired, retry with the same ids and never resubmit the paid create. A walk-through of the batch form is in parallel video jobs with wait any/all.
What does queued mean for the user?
Queued jobs are accepted but not processing; workers move them to processing under the per-workspace guard. A subagent should report queued as waiting, not failed, and treat 429 queue_full as the signal to stop submitting.
Sources
Related posts
More in Developers
- Music API 400 model_not_found: fix an unknown model id
An unknown model on POST /v1/music-router/generate fails with 400 model_not_found and a catalog_url. Use an id from GET /v1/music-router/models or omit model.
- Sume music job says sume/music-auto: which engine ran?
job.model echoes the id you requested; job.request.routed_model names the engine that ran, such as lyria-3.5. Read both fields on the job envelope.
- n8n Wait node under 65 seconds: how to poll a Sume job
n8n keeps waits under 65 seconds in memory and saves longer ones to the database. Poll a Sume job with next_poll_after_seconds, or switch to a webhook resume.
- Next.js image SSRF fix: narrow remotePatterns for Sume
Next.js patched an Image Optimization SSRF via allow-listed remote URLs. Allow only media.sume.com for Sume images, and never raw provider hosts.
Written by Sume