Runway enterprise exception requests vs Sume's queue-first admission

Runway's docs mention enterprise exception requests for higher volume. Sume takes extra jobs as queued and sets concurrency by plan. How to plan a big batch.

4 min readSume
All posts

Runway's developer docs list enterprise exception requests for higher volume, while Sume sets generation concurrency by plan and accepts extra valid jobs as queued. For a large batch the practical difference is that on Sume you will see the limit in your own submit responses and can plan around it without asking anyone first.

The Runway developer docs also list Seedance 2.5 at 1080p and up to 30 seconds, Gen 4.5, Aleph 2.0 and GPT Image 2. Sume lists Seedance 2.5 in its Video Router catalog too, with its own limits.

What the Runway page lists

As listed on the page.

Runway developer docs (read 2026-10-03)
ItemWhat the page says
Seedance 2.51080p, up to 30 seconds
Other modelsGen 4.5, Aleph 2.0, GPT Image 2
Higher volumeEnterprise exception requests

How Sume sets limits

Generation concurrency on Sume is plan-only. Prepaid top-ups do not raise it; admin overrides can raise the effective limit. Queue capacity defaults to the larger of 3 and five times the concurrency limit. The dashboard Concurrency tab is the source of truth, exposed as generation_limits.concurrency_limit.

Sume plan defaults, from the generation admission docs (read 2026-10-03)
PlanProcessing concurrencyQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

What queue-first means in practice

Concurrency is a dispatch limit, not a submit limit. With a limit of 1 you can submit several valid jobs at once and Sume can return all of them as queued; one moves to processing at a time. Do not treat queued as a failure. When accepted capacity is full, the next submit gets 429 queue_full, which is different from 429 rate_limited, the request-volume limit.

Enterprise defaults to 20 and uses admin overrides for higher contract limits, so a higher ceiling is a conversation about your workspace, not a form you must file before you start.

Planning a batch

A simple plan keeps you inside the limits without guesswork.

  • Read generation_limits from the first submit and use the effective concurrency_limit, not the table above.
  • Compute in-flight room as the concurrency limit minus active and queued jobs, floored at zero.
  • Submit in waves, and use queue_capacity_remaining to decide when to stop adding work.
  • Give every submit an Idempotency-Key built from your own item id, so a retry returns the original job.

The Seedance 2.5 limits on Sume

On the Video Router, seedance-2.5 accepts 4 to 30 seconds at 480p, 720p or 1080p. Billing is provider list times 1.25 on every model. The docs advise reading capabilities from GET /v1/video-router/models before assuming a limit, since each model has its own envelope.

The same model id on another service may have different limits and different pricing, so compare duration, resolution and per-second price for the specific model you plan to run, not the family name.

When you hit queue_full

The documented response is short.

On 429 queue_full:
  stop adding generation work for the workspace
  poll existing jobs until at least one is terminal
  cancel queued jobs you no longer need
  retry with the same Idempotency-Key after capacity opens
  honor retry-after when present

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume