Check GET /v1/balance before a Sume bulk run: failures land per item

A bulk create returns 202 even if the wallet cannot fund every child; unfunded items fail one by one. Compare GET /v1/balance with your spend caps first.

5 min readSume
All posts

Read GET /v1/balance before you queue a bulk run and compare it with the worst case your items can spend. A bulk create answers 202 once the queue exists, but child runs still go through ordinary admission, so one the wallet cannot fund fails as that item, after the 202 has already come back.

The relevant sentence is on the Bulk runs page: a child that fails to start is that item failed; the queue create has already returned 202. A preflight turns a mid-batch surprise into an early stop.

What does GET /v1/balance return?

It returns the workspace's USD balance for the authenticated key, as data.balance. A missing balance row is an explicit zero rather than an error. The fields that matter here are available_amount_usd_micros, available_amount_usd_cents and state, which is funded or empty.

Amounts are in micros, so 1,000,000 micros is one dollar. Do the arithmetic in integers and convert only for display.

What is the worst case to compare against?

Every Format has a generation spend cap, and a run can never spend past its effective cap. A request can set generation_spend_cap_usd for the run, which must be above 0 and at most 500. A Format that never named a cap reports the platform default of $400, so an item with no explicit cap makes the worst case huge and the preflight useless.

So set a cap on every item. Then the ceiling for the batch is the sum of those caps, and the check is simply: is the available balance at least that sum? That is conservative, since most runs spend less, but it is the number the platform will not let you exceed.

What does the check look like?

This reads the balance, sums per-item caps in micros, and refuses to queue if the ceiling exceeds what is available. It is advisory: another process can spend between the read and the create, so keep handling per-item failures regardless.

const rows = [{ id: 1, url: "https://example.com/1.jpg" }]; // your batch
const H = { "x-api-key": process.env.SUME_API_KEY! };
const items = rows.map((r) => ({
  instruction: `clip ${r.id}`,
  input: { url: r.url },
  generation_spend_cap_usd: 3,
}));

const bal = await fetch("https://api.sume.com/v1/balance", { headers: H })
  .then((r) => r.json());
const available = bal.data.balance.available_amount_usd_micros;
const ceiling = items.reduce((s, i) => s + i.generation_spend_cap_usd * 1_000_000, 0);

if (bal.data.balance.state === "empty" || ceiling > available) {
  throw new Error(`worst case ${ceiling} > available ${available} micros`);
}
const res = await fetch("https://api.sume.com/v1/formats/acme/product-promo/bulk-runs", {
  method: "POST",
  headers: { ...H, "Content-Type": "application/json", "Idempotency-Key": "batch-2026-10-02" },
  body: JSON.stringify({ concurrency: 3, items }),
});
console.log(res.status); // 202

What does the preflight not cover?

It does not replace handling failures. A bulk queue takes up to 100 items with a concurrency window of 1 to 16, and admission also checks workspace generation concurrency and each item's cap. Poll GET /v1/format-run-queues/{queue_id} for per-item status and read the failed ones.

A funded wallet does not guarantee every child starts: admission also checks workspace generation concurrency and each item's cap, and a 402 insufficient_credits has next_action of add_funds, so retrying without topping up returns the same answer.

Finally, balance is also visible through the CLI as sume balance, and GET /v1/usage lists ledger entries such as reservations, captures and refunds if you want to reconcile after the batch.

How should I read the result afterwards?

Poll the queue and treat each item separately. GET /v1/format-run-queues/{queue_id} returns counts plus per-item status, and each child is an ordinary run you can read at GET /v1/format-runs/{run_id}. The queue has no webhook of its own; communication.webhook_url is set per item, and a run webhook is delivered only for completed or failed children.

There is no public list-queues or cancel-queue endpoint. To stop work, cancel each child with POST /v1/format-runs/{run_id}/cancel, which is idempotent. So the cheapest control you have is the one before the create: the preflight, a sane concurrency and an explicit cap on every item.

After the batch, reconcile. GET /v1/usage accepts a run_id filter that sums a Format run's generation jobs, and its summary carries the cap accounting and the debited amount, so you can compare what you reserved with what was spent.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume