p-limit npm: cap concurrent AI API jobs in Node.js

p-limit runs at most n promise-returning functions at once. For paid AI jobs, wrap the submit and the wait, not just the POST, and set n to your limit.

5 min readSume
All posts

p-limit is an npm package that runs promise-returning and async functions with limited concurrency: pLimit(n) returns a limit function, and at most n of the functions you pass to it run at once while the rest wait their turn. To cap concurrent AI API jobs, pass limit a function that submits the job and waits for its final status, not one that only sends the POST.

The p-limit facts below come from its README, read 2026-09-28. The job API in the example is Sume's; its limits come from the generation admission docs.

How do I use p-limit?

Install it with npm install p-limit; the README imports it as import pLimit from 'p-limit' and says it works in Node.js and browsers. Its whole API fits in one table.

From the p-limit README, read 2026-09-28.
MemberWhat it does
pLimit(concurrency)Returns a limit function; concurrency is a number (minimum 1) or { concurrency }
limit(fn, ...args)Returns the promise from calling fn(...args) once a slot is free
limit.map(iterable, mapper)Runs mapper(item, index) over the inputs; equivalent to Promise.all over limit calls
limit.activeCount / limit.pendingCountPromises running now / waiting to start
limit.clearQueue()Discards calls that have not started; does not cancel running ones
rejectOnClear optionRejects cleared calls with an AbortError instead of leaving them unresolved (default false)

How do I limit concurrent AI video jobs with p-limit?

Wrap the submit and the polling loop in one function and hand it to limit.map. POST /v1/videos returns a job with a polling_url at once; the loop polls it until the status is final (Video generation). Each item gets its own stable Idempotency-Key, so rerunning the script with the same prompts replays accepted jobs instead of paying twice. Run this on a server: the key belongs in a server environment variable (Authentication).

import pLimit from "p-limit";

const API = "https://api.sume.com/v1/videos";
const auth = { Authorization: `Bearer ${process.env.SUME_API_KEY}` };
const DONE = ["completed", "failed", "cancelled"];
const limit = pLimit(4); // your workspace's concurrency_limit
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

async function makeVideo(prompt: string, i: number) {
  const res = await fetch(API, {
    method: "POST",
    headers: { ...auth, "Content-Type": "application/json", "Idempotency-Key": `batch-7-${i}` },
    body: JSON.stringify({ model: "seedance-2", prompt, resolution: "720p" }),
  });
  if (!res.ok) return { error: res.status, body: await res.json() };
  let job = await res.json();
  while (!DONE.includes(job.status)) {
    await sleep(30_000);
    job = await (await fetch(job.polling_url, { headers: auth })).json();
  }
  return job;
}

const results = await limit.map(prompts, makeVideo);

Why wrap the whole job instead of the fetch?

A submit returns in a moment, but the job runs for minutes. Limit only the fetch and every prompt is submitted almost at once. Sume accepts jobs beyond your processing limit as queued, reserves each one's estimated cost at submit, and refuses new ones with 429 queue_full once the queue is also full (Generation admission). Wrapping submit plus wait keeps in-flight work at n. asyncio Semaphore is the same pattern in Python.

Don't call limit again inside makeVideo: the README warns that calling the same limit function inside a function it already limits can deadlock. Use a separate limiter for inner work.

What should n be?

Start from your workspace's effective concurrency_limit; the docs call the dashboard Concurrency tab its source of truth (Generation admission). Video job concurrency and queueing shows how to size from a live generation_limits snapshot when other processes share the workspace.

A limit function only counts the calls made through it, in one process. Two workers that each create pLimit(4) can have eight jobs in flight against the same workspace, so split the limit between them.

Polling spends a read budget separate from the write budget, so the status loop does not eat into submits (Authentication).

What happens when one job fails?

limit.map behaves like Promise.all, so one thrown error rejects the whole batch while the other jobs keep running and spending. That's why the example returns an error object instead of throwing.

  • 402 insufficient_credits: stop the batch. Call limit.clearQueue() so nothing new starts, with rejectOnClear if you await the results.
  • 429 queue_full: wait for a running job to finish, then submit again under a new key. In current code a same-key resend returns the stored queue_full.
  • 429 rate_limited: wait for retry-after, then retry with the same key.
  • A local timeout does not cancel a Sume job, so never resubmit just because your loop gave up; keep polling the stored job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume