Runway: no RPM limit but a daily cap; Sume answers 429 two ways
Runway documents no per-minute limit but a rolling 24-hour generation cap. Sume returns 429 queue_full or rate_limited, with ratelimit headers.

Runway limits how many generations you start in a rolling 24 hours, not how many requests you send per minute. Sume has the opposite shape: it applies request rate limits and a queue, and it answers 429 with one of two codes, queue_full or rate_limited, which need different responses.
Here are the limits as the two pages document them, and a small client that tells the 429s apart.
Runway's limits
| Item | Value |
|---|---|
| Requests per minute | No RPM limit |
| Daily generation cap | Rolling 24 hour window; 50-200 at Tier 1 |
| Throttled tasks | Queued in submission order |
| Concurrency | 1-2 at Tier 1, rising to 20 at Tier 5 |
Sume's limits
- Rate-limit headers on public responses can include ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after.
- None of the four admission controls on Sume's page is a daily generation count.
| Code | Meaning | What to do |
|---|---|---|
| 429 queue_full | No accepted generation capacity left in the workspace | Wait for jobs to finish or cancel queued ones, then retry with the same idempotency key |
| 429 rate_limited | Request volume exceeded an abuse-protection limit | Back off using retry-after when present |
| 402 insufficient_credits | Balance cannot cover the estimate | Change plan, wait for Gen$, or submit something cheaper |
| 503 provider_capacity_exceeded | Dispatch queue is full | Retry later with the same idempotency key |
A client that separates the two
Retry only when the 429 is rate_limited, and stop on queue_full.
async function submit(body, key) {
for (let attempt = 0; attempt < 5; attempt++) {
const res = await fetch("https://api.sume.com/v1/image-1.0/generate", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify(body),
});
if (res.status !== 429) return res;
const { error } = await res.json();
if (error.code === "queue_full") throw new Error("queue full");
const wait = Number(res.headers.get("retry-after")) || 2 ** attempt;
await new Promise((r) => setTimeout(r, wait * 1000));
}
throw new Error("still rate limited");
}
async function main() {
const res = await submit({ prompt: "Matte black bottle on marble", mode: "async" }, "hero-001");
console.log(res.status);
}
main();Planning
For Runway, track generations over a rolling day rather than a request rate. For Sume, size in-flight work from generation_limits (concurrency_limit, queued_generation_jobs and queue_capacity_remaining) instead of firing a burst, and keep one idempotency key per intended job. See the Sume generation admission docs.
Sources
Related posts
More in Comparisons
- Runway SDK needs Node 18 or Python 3.8; Sume SDK is Node 18, no deps
Runway lists Node 18+ and Python 3.8+. Sume's TypeScript SDK 0.2.0 runs on Node 18+, Bun, Deno or Workers, no runtime deps; no Python SDK is documented.
- Throttled Runway tasks queue in order; Sume shows queued
Runway queues throttled API tasks in submission order. Sume accepts valid jobs as queued: Free 5, Pro 20, Startup 40, Scale 100 slots, then 429 queue_full.
- Runway WAN3 vs WAN3 Prime: Prime is 36 to 40% more; Sume lists one Wan
Runway prices WAN3 at 5, 10, 20 credits per second and WAN3 Prime at 6.8, 14, 28: 36% to 40% more. Sume lists one wan-3.0, at $0.0625, $0.125, $0.25.
- Seedance 2.5 takes 10 video references; Omni 1.1 limits them to 3 s
Seed lists up to 10 video clips per Seedance 2.5 pass. Google's Omni 1.1 Flash video references are up to 3 s. What that means for source clips on Sume.
Written by Sume