Throttled Runway tasks queue in order; Sume shows queued

Runway queues throttled API tasks in submission order. Sume accepts valid jobs as queued: Free 5, Pro 20, Startup 40, Scale 100 slots, then 429 queue_full.

4 min readSume
All posts

Runway says tasks throttled by your tier are queued in submission order. Sume accepts valid paid jobs as queued while queue capacity remains, and moves them to processing as concurrency slots open. The difference is in what you can read off the system: Sume publishes the queue size for each plan and an explicit error when it is full.

On Sume, do not treat queued as a failure: store the job id and poll.

Runway: concurrency by tier, queue in submission order

Runway's usage-tier page lists five tiers. Concurrency is 1-2 at tier 1, then 3, 5, 10 and 20. The page also states that there is no requests-per-minute limit, and that tasks above your concurrency are throttled and queued in submission order. It does not publish a queue length in the text I read, so I make no claim about one.

Runway tier concurrency, read 2026-10-05
TierConcurrency
11-2
23
35
410
520

Sume: concurrency plus a published queue

On Sume, concurrency is a dispatch limit and not a submit limit. A workspace at its processing cap can still accept new jobs as queued. Queue capacity is the default max(3, concurrency_limit x 5), and accepted job capacity is concurrency_limit + queue. The docs table below gives the plan defaults. The effective values for your workspace are in generation_limits on each submit response, which the docs say to prefer over the static table.

Sume plan defaults, docs read 2026-10-05
PlanProcessingQueueAccepted (processing + queued)
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

Arithmetic check and what happens at the edge

The accepted column is the sum of the other two: 1 + 5 = 6, 4 + 20 = 24, 8 + 40 = 48, 20 + 100 = 120. On a Pro workspace you can therefore have 24 paid jobs in flight at once, 4 running and 20 waiting. The 25th valid submit fails with 429 queue_full, which is different from 429 rate_limited (request volume) and 402 insufficient_credits (balance).

When you see queue_full, the docs say to stop adding work, wait for a job to reach a terminal state or cancel queued jobs you no longer need, and retry with the same idempotency key. Cancel works only before generation starts. After that the API answers 409 job_generation_already_started.

What to do in a client

  • Treat queued and processing as normal non-terminal states and poll with backoff until terminal is true.
  • Read queue_capacity_remaining from generation_limits before a burst. The docs size new in-flight work as max(0, concurrency_limit - active - queued), capped at the remaining queue capacity.
  • Do not submit the paid request again because a local worker timed out. Retry with the same Idempotency-Key.
  • Sume does not show a per-job queue position or ETA today, so do not build a countdown on it.

Worked example: sizing a launch batch

Say a Pro workspace wants to render 30 product clips. With 4 processing and 20 queued slots, 24 jobs are accepted and the other 6 would be refused with queue_full. Submit in waves instead: the docs give wave_size_hint as max(1, floor(queue_capacity_remaining x 0.75)). On an empty Pro workspace, queue_capacity_remaining is 24, so 24 x 0.75 = 18, which matches the 18 in the docs example. Submit 18, poll until some finish, read the fresh snapshot and submit more.

The docs warn that the hint is a submission-wave hint only. It is not a concurrency limit and it should never be shown as one. On the Runway side, the equivalent step is to keep your in-flight count under your tier's concurrency and let throttled tasks wait their turn.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume