Pace bulk Sume submits with generation_limits, not wave_size_hint

Size in-flight work as concurrency minus active minus queued, capped by queue capacity. A short Node pacer that reads generation_limits from each submit.

5 min readSume
All posts

To submit a large batch to Sume without hitting queue_full, read data.generation_limits from each submit response and allow new work only up to min(max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), queue_capacity_remaining). Do not use wave_size_hint for this: the docs call it a submission-wave hint, not a concurrency limit, and say never to use it to size in-flight work.

The difference is large. The docs' own example has a concurrency limit of 100, a queue of 500 and nothing running: queue_capacity_remaining is 600 and wave_size_hint is 450, but the new in-flight budget is 100. With 30 processing and 10 queued the budget is 60.

Fields used

Use the effective concurrency_limit, not plan_concurrency_limit, because an admin override can change it. The counts are a snapshot and can be stale a moment after the response.

generation_limits fields, read 2026-10-02 from docs.sume.com
FieldMeaning
concurrency_limitEffective maximum same-workspace jobs in processing
active_generation_jobsJobs currently processing
queued_generation_jobsJobs currently queued
queue_capacity_remainingQueue budget plus idle processing seats before queue_full
wave_size_hintmax(1, floor(remaining x 0.75)); a hint only

The pacer

submit(item) should send one create with an Idempotency-Key and return the job id plus the generation_limits object. waitForAny(ids) is your status poll and returns one id that reached a terminal state. When the budget hits zero the loop waits for a finish and then submits once, which brings back a fresh snapshot.

// New in-flight budget, per the generation admission docs. wave_size_hint is NOT this.
export const headroom = (g) =>
  Math.min(
    Math.max(0, g.concurrency_limit - g.active_generation_jobs - g.queued_generation_jobs),
    g.queue_capacity_remaining,
  );

// submit(item) -> { jobId, limits }   limits = data.generation_limits from the submit response
// waitForAny(ids) -> id of a job that reached a terminal state
export async function runPaced(items, submit, waitForAny) {
  const pending = [...items];
  const running = new Set();
  let budget = 1; // the first submit returns the first real snapshot
  while (pending.length || running.size) {
    if (pending.length && budget > 0) {
      const { jobId, limits } = await submit(pending.shift());
      running.add(jobId);
      if (limits) budget = headroom(limits); // fresh snapshot replaces the guess
      else budget = 0; // no counts: wait for a finish, then refresh
    } else {
      running.delete(await waitForAny([...running]));
      budget = 1; // a slot freed; the next submit refreshes the snapshot
    }
  }
}

Why not just submit everything

Sume accepts valid jobs as queued while queue capacity remains, so over-submitting works until it does not. The docs describe client pacing as keeping outstanding work within the processing cap, while the API separately keeps accepting queued work. Pacing also leaves room for other callers on the workspace and keeps cancel simple: fewer queued jobs to clean up.

Limits

The pacer is a sketch, not a tested client. Live counts include other callers, so another service on the same workspace can still fill the queue and return 429 queue_full. In that case stop adding work, let existing jobs finish, and retry with the same idempotency key. A tight loop on a limits-less response falls back to one job at a time.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume