Shared Format run stuck queued: whose concurrency limit applies?
A partner's run on your shared Format uses the partner's concurrency slot, so the partner's plan limit and queue decide when it starts, not yours.

When a partner's run on your shared Format sits queued, look at the partner's workspace, not yours. A shared run uses the grantee's concurrency slot, so the grantee's plan limit, active jobs and queue capacity decide when it starts.
Why it is the grantee's limit
The Formats docs state that the run, its spend, its concurrency slot and its media belong to the grantee's workspace. Generation concurrency itself is plan-only: prepaid top-ups do not raise it, and an admin override can. The dashboard Concurrency tab is the source of truth for a workspace, and the API shows the effective figure as generation_limits.concurrency_limit.
The default plan table in the admission docs is a static reference; always prefer the effective field.
| Plan | Processing concurrency | Queue capacity (default) | Accepted job capacity |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
Org floor and overrides
Org workspaces have a floor of 10, and Enterprise defaults to 20 with admin overrides for contract limits. A partner on a Free plan therefore runs one generation job at a time on your Format, however large your own limit is.
Concurrency is a dispatch limit, not a submit limit. A partner at its limit can still submit and get queued, until the queue is full and the answer is 429 queue_full.
What each side can do
The owner cannot raise a partner's concurrency; the owner can only keep the Format lean so each run needs fewer jobs. The partner can read generation_limits from its own submit responses and pace new work to the headroom: concurrency limit minus active minus queued, capped at the remaining queue capacity.
Do not tell a partner that a queued run is a failure. It is a normal state with no position or ETA, and it moves to processing when a slot opens.
def headroom(gl):
room = max(0, gl["concurrency_limit"]
- gl["active_generation_jobs"]
- gl["queued_generation_jobs"])
return min(room, gl["queue_capacity_remaining"])How a partner can confirm it
The partner reads generation_limits from any generation submit response in its own workspace. If active_generation_jobs equals concurrency_limit, the run is waiting for a slot of theirs, and the cure is on their side: wait, cancel queued work they no longer need, or move to a plan with a higher limit. Cancellation works only before generation starts.
If queue_capacity_remaining is zero, their next submit fails with queue_full. Ask them to stop adding work until a job finishes, then retry with the same idempotency key.
Both of these are visible without any help from you, which is why a support request about a slow shared run should start with the grantee's counts.
Limits
Your own jobs and the partner's jobs do not compete for the same slots, since each workspace has its own. That is the practical benefit of a grant over lending your team key: lending a key puts the partner's work in your concurrency window and on your bill.
Sources
Related posts
More in Formats
- Sume bulk Format runs: 100 items, one concurrency window, cost control
POST /v1/formats/{handle}/{slug}/bulk-runs queues up to 100 ordinary Format runs. Set concurrency and a per-run cap, and poll the queue with format-run-queues.
- Sume catalog Formats for ads: which of 27 slugs to read first
The Sume Format catalog lists 27 slugs. Group them by the ad job their names suggest, then confirm each with GET /v1/formats/sume/{slug} before you call it.
- Format run input extra keys: context only, not echoed in output
Keys your Format does not read, like a sheet row number, reach the run as context and do not return in output. Store them beside the 202's id and thread id.
- Format status_url never holds output: poll it, then fetch the result
A Sume Format run's status_url returns a small poll payload with no output or artifacts. Poll it with backoff up to expires_at, then call result_url once.
Written by Sume