Capacity fallback vs Sume queued jobs at full concurrency
Runway's Model Router can fall back to another model at its concurrency limit. Sume instead accepts the job as queued on the same model until a slot opens.

When a Sume workspace is at its generation concurrency limit, a new valid job is accepted as queued and waits for a slot on the model you asked for. It is not rerouted to another model. Runway's Model Router, per its Jul 30, 2026 changelog, can instead fall back to the next-best eligible model when you hit your concurrency limit, if you opt in per router.
Vendor text is from the Runway changelog as read 2026-10-01. Sume behavior is from Generation admission.
What does Runway's capacity fallback do?
The changelog says that at the concurrency limit on a model, Model Router can automatically fall back to the next-best eligible model with capacity instead of queuing the request. It is opt-in per model router in the dev portal, and it is described as a way to scale past per-model concurrency limits.
What does Sume do at the limit?
Sume's docs say jobs may start immediately or wait in queued until a workspace concurrency slot opens. Concurrency is a dispatch limit, not a submit limit, and the same job later moves to processing. The model never changes after submit.
| Question | Runway Model Router | Sume |
|---|---|---|
| Request at the limit | Can fall back to another eligible model (opt-in) | Accepted as queued on the same model |
| Model changes? | Yes, if fallback is enabled | No |
| What raises the limit | Not stated in the entry | Plan only; prepaid top-ups do not raise it |
| When the queue is full | Not stated in the entry | New paid submissions fail with 429 queue_full |
How do I see how much room I have?
The admission docs list the fields concurrency_limit, queued_jobs_limit, accepted_generation_jobs_limit, active_generation_jobs, and queued_generation_jobs. Queue capacity defaults to max(3, concurrency_limit × 5). For the error path, see queue_full vs concurrency full.
Can I get a different model when the queue is long?
Only by choosing it yourself. You name a model per request, or pass sume/auto, which the video docs describe as letting Sume pick the family up front. Neither swaps a job that is already queued. If your workflow needs a fallback, submit to a second model explicitly from your own code.
Sources
Related posts
More in Developers
- Muse Voice Transcribe 16 kHz mono PCM: ffmpeg vs Sume audio-detach
Muse Voice Transcribe wants mono 16-bit PCM at 24 or 16 kHz. On Sume, audio-detach with 16000 Hz and mono makes the STT input wav from a hosted video.
- Muse Voice Transcribe WebSocket streaming vs Sume STT job waits
Muse Voice Transcribe streams audio over one WebSocket and returns cumulative partials. Sume STT is a job: wait up to 30 seconds, then poll status_url.
- Nano Banana batch API: Gemini's 24 h batch vs Sume async jobs
Gemini's Batch API trades up to 24 hours of turnaround for higher rate limits. Sume has no batch tier for images: send async or webhook jobs per request.
- Nano Banana Pro 21:9: Gemini's ratio list vs Sume's catalog
Gemini's image docs list 21:9 among ten ratios. Sume's Nano Banana Pro and Nano Banana 2 catalogs include 21:9 too; GPT Image 2.5's list does not.
Written by Sume