Capacity fallback vs Sume queued jobs at full concurrency

Runway's Model Router can fall back to another model at its concurrency limit. Sume instead accepts the job as queued on the same model until a slot opens.

4 min readSume
All posts

When a Sume workspace is at its generation concurrency limit, a new valid job is accepted as queued and waits for a slot on the model you asked for. It is not rerouted to another model. Runway's Model Router, per its Jul 30, 2026 changelog, can instead fall back to the next-best eligible model when you hit your concurrency limit, if you opt in per router.

Vendor text is from the Runway changelog as read 2026-10-01. Sume behavior is from Generation admission.

What does Runway's capacity fallback do?

The changelog says that at the concurrency limit on a model, Model Router can automatically fall back to the next-best eligible model with capacity instead of queuing the request. It is opt-in per model router in the dev portal, and it is described as a way to scale past per-model concurrency limits.

What does Sume do at the limit?

Sume's docs say jobs may start immediately or wait in queued until a workspace concurrency slot opens. Concurrency is a dispatch limit, not a submit limit, and the same job later moves to processing. The model never changes after submit.

Behavior at full concurrency, from the Runway changelog and Sume docs, read 2026-10-01.
QuestionRunway Model RouterSume
Request at the limitCan fall back to another eligible model (opt-in)Accepted as queued on the same model
Model changes?Yes, if fallback is enabledNo
What raises the limitNot stated in the entryPlan only; prepaid top-ups do not raise it
When the queue is fullNot stated in the entryNew paid submissions fail with 429 queue_full

How do I see how much room I have?

The admission docs list the fields concurrency_limit, queued_jobs_limit, accepted_generation_jobs_limit, active_generation_jobs, and queued_generation_jobs. Queue capacity defaults to max(3, concurrency_limit × 5). For the error path, see queue_full vs concurrency full.

Can I get a different model when the queue is long?

Only by choosing it yourself. You name a model per request, or pass sume/auto, which the video docs describe as letting Sume pick the family up front. Neither swaps a job that is already queued. If your workflow needs a fallback, submit to a second model explicitly from your own code.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume