Shared Format run stuck queued: whose concurrency limit applies?

A partner's run on your shared Format uses the partner's concurrency slot, so the partner's plan limit and queue decide when it starts, not yours.

4 min readSume
All posts

When a partner's run on your shared Format sits queued, look at the partner's workspace, not yours. A shared run uses the grantee's concurrency slot, so the grantee's plan limit, active jobs and queue capacity decide when it starts.

Why it is the grantee's limit

The Formats docs state that the run, its spend, its concurrency slot and its media belong to the grantee's workspace. Generation concurrency itself is plan-only: prepaid top-ups do not raise it, and an admin override can. The dashboard Concurrency tab is the source of truth for a workspace, and the API shows the effective figure as generation_limits.concurrency_limit.

The default plan table in the admission docs is a static reference; always prefer the effective field.

Default processing concurrency and queue by plan, from Generation admission (read 2026-10-05)
PlanProcessing concurrencyQueue capacity (default)Accepted job capacity
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

Org floor and overrides

Org workspaces have a floor of 10, and Enterprise defaults to 20 with admin overrides for contract limits. A partner on a Free plan therefore runs one generation job at a time on your Format, however large your own limit is.

Concurrency is a dispatch limit, not a submit limit. A partner at its limit can still submit and get queued, until the queue is full and the answer is 429 queue_full.

What each side can do

The owner cannot raise a partner's concurrency; the owner can only keep the Format lean so each run needs fewer jobs. The partner can read generation_limits from its own submit responses and pace new work to the headroom: concurrency limit minus active minus queued, capped at the remaining queue capacity.

Do not tell a partner that a queued run is a failure. It is a normal state with no position or ETA, and it moves to processing when a slot opens.

def headroom(gl):
    room = max(0, gl["concurrency_limit"]
               - gl["active_generation_jobs"]
               - gl["queued_generation_jobs"])
    return min(room, gl["queue_capacity_remaining"])

How a partner can confirm it

The partner reads generation_limits from any generation submit response in its own workspace. If active_generation_jobs equals concurrency_limit, the run is waiting for a slot of theirs, and the cure is on their side: wait, cancel queued work they no longer need, or move to a plan with a higher limit. Cancellation works only before generation starts.

If queue_capacity_remaining is zero, their next submit fails with queue_full. Ask them to stop adding work until a job finishes, then retry with the same idempotency key.

Both of these are visible without any help from you, which is why a support request about a slow shared run should start with the grantee's counts.

Limits

Your own jobs and the partner's jobs do not compete for the same slots, since each workspace has its own. That is the practical benefit of a grant over lending your team key: lending a key puts the partner's work in your concurrency window and on your bill.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume