Runway: no RPM limit but a daily cap; Sume answers 429 two ways

Runway documents no per-minute limit but a rolling 24-hour generation cap. Sume returns 429 queue_full or rate_limited, with ratelimit headers.

4 min readSume
All posts

Runway limits how many generations you start in a rolling 24 hours, not how many requests you send per minute. Sume has the opposite shape: it applies request rate limits and a queue, and it answers 429 with one of two codes, queue_full or rate_limited, which need different responses.

Here are the limits as the two pages document them, and a small client that tells the 429s apart.

Runway's limits

Runway API limits, read 2026-10-05
ItemValue
Requests per minuteNo RPM limit
Daily generation capRolling 24 hour window; 50-200 at Tier 1
Throttled tasksQueued in submission order
Concurrency1-2 at Tier 1, rising to 20 at Tier 5

Sume's limits

  • Rate-limit headers on public responses can include ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after.
  • None of the four admission controls on Sume's page is a daily generation count.
Sume 429 responses, Sume docs checked 2026-10-05
CodeMeaningWhat to do
429 queue_fullNo accepted generation capacity left in the workspaceWait for jobs to finish or cancel queued ones, then retry with the same idempotency key
429 rate_limitedRequest volume exceeded an abuse-protection limitBack off using retry-after when present
402 insufficient_creditsBalance cannot cover the estimateChange plan, wait for Gen$, or submit something cheaper
503 provider_capacity_exceededDispatch queue is fullRetry later with the same idempotency key

A client that separates the two

Retry only when the 429 is rate_limited, and stop on queue_full.

async function submit(body, key) {
  for (let attempt = 0; attempt < 5; attempt++) {
    const res = await fetch("https://api.sume.com/v1/image-1.0/generate", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${process.env.SUME_API_KEY}`,
        "Content-Type": "application/json",
        "Idempotency-Key": key,
      },
      body: JSON.stringify(body),
    });
    if (res.status !== 429) return res;
    const { error } = await res.json();
    if (error.code === "queue_full") throw new Error("queue full");
    const wait = Number(res.headers.get("retry-after")) || 2 ** attempt;
    await new Promise((r) => setTimeout(r, wait * 1000));
  }
  throw new Error("still rate limited");
}

async function main() {
  const res = await submit({ prompt: "Matte black bottle on marble", mode: "async" }, "hero-001");
  console.log(res.status);
}
main();

Planning

For Runway, track generations over a rolling day rather than a request rate. For Sume, size in-flight work from generation_limits (concurrency_limit, queued_generation_jobs and queue_capacity_remaining) instead of firing a burst, and keep one idempotency key per intended job. See the Sume generation admission docs.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume