Runway enterprise exception requests vs Sume's queue-first admission
Runway's docs mention enterprise exception requests for higher volume. Sume takes extra jobs as queued and sets concurrency by plan. How to plan a big batch.

Runway's developer docs list enterprise exception requests for higher volume, while Sume sets generation concurrency by plan and accepts extra valid jobs as queued. For a large batch the practical difference is that on Sume you will see the limit in your own submit responses and can plan around it without asking anyone first.
The Runway developer docs also list Seedance 2.5 at 1080p and up to 30 seconds, Gen 4.5, Aleph 2.0 and GPT Image 2. Sume lists Seedance 2.5 in its Video Router catalog too, with its own limits.
What the Runway page lists
As listed on the page.
| Item | What the page says |
|---|---|
| Seedance 2.5 | 1080p, up to 30 seconds |
| Other models | Gen 4.5, Aleph 2.0, GPT Image 2 |
| Higher volume | Enterprise exception requests |
How Sume sets limits
Generation concurrency on Sume is plan-only. Prepaid top-ups do not raise it; admin overrides can raise the effective limit. Queue capacity defaults to the larger of 3 and five times the concurrency limit. The dashboard Concurrency tab is the source of truth, exposed as generation_limits.concurrency_limit.
| Plan | Processing concurrency | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
What queue-first means in practice
Concurrency is a dispatch limit, not a submit limit. With a limit of 1 you can submit several valid jobs at once and Sume can return all of them as queued; one moves to processing at a time. Do not treat queued as a failure. When accepted capacity is full, the next submit gets 429 queue_full, which is different from 429 rate_limited, the request-volume limit.
Enterprise defaults to 20 and uses admin overrides for higher contract limits, so a higher ceiling is a conversation about your workspace, not a form you must file before you start.
Planning a batch
A simple plan keeps you inside the limits without guesswork.
- Read
generation_limitsfrom the first submit and use the effectiveconcurrency_limit, not the table above. - Compute in-flight room as the concurrency limit minus active and queued jobs, floored at zero.
- Submit in waves, and use
queue_capacity_remainingto decide when to stop adding work. - Give every submit an
Idempotency-Keybuilt from your own item id, so a retry returns the original job.
The Seedance 2.5 limits on Sume
On the Video Router, seedance-2.5 accepts 4 to 30 seconds at 480p, 720p or 1080p. Billing is provider list times 1.25 on every model. The docs advise reading capabilities from GET /v1/video-router/models before assuming a limit, since each model has its own envelope.
The same model id on another service may have different limits and different pricing, so compare duration, resolution and per-second price for the specific model you plan to run, not the family name.
When you hit queue_full
The documented response is short.
On 429 queue_full:
stop adding generation work for the workspace
poll existing jobs until at least one is terminal
cancel queued jobs you no longer need
retry with the same Idempotency-Key after capacity opens
honor retry-after when presentSources
Related posts
More in Comparisons
- Seedance 2.5, Wan 3.0 and Omni Flash: where each one is sold
Runway lists seedance2_5, wan3 and gemini_omni_flash_1.1 as API ids, Luma hosts Seedance 2.5, and Sume routes all three. How to compare.
- Seedance 2.5: fal's 30 s at 720p listing vs Sume's catalog limits
fal.ai/models says Seedance 2.5 is a native 30-second clip at up to 720p. Sume's catalog lists 4 to 30 s at 480p, 720p and 1080p. Read the live values.
- Stable Diffusion alternatives in 2026: open weights or hosted API
SD 3.5 is still Stability's newest flagship image model. FLUX.2 [klein], Qwen-Image 2.0 and Z-Image Turbo are the open options. How to pick between them.
- Suno Speech beta: voice and music in one pass, or separate tracks?
Suno's Speech beta makes voice and music in one track. Its blog lists wandering accents and long pauses. When to prefer separate TTS, music and a timeline mix.
Written by Sume