Format run cap headroom: usage.cap limit, counted and remaining

usage.cap on a run receipt splits the spend-cap check into limit, counted and remaining USD micros. Read it to see how close a run is to failing.

5 min readSume
All posts

To see how near a Format run is to its spend cap, read usage.cap on the receipt. It gives the cap check in three parts: limit_usd_micros, counted_usd_micros and remaining_usd_micros. counted_usd_micros is the same number as usage.billable_amount_usd_micros. When the counted amount reaches the limit, the run stops spending and ends as failed, and usage shows how near the cap it got.

What each number is

Sume enforces the cap of a run against billable_amount_usd_micros. That total counts reserved and captured generation amounts, rises while the run is in progress and settles when the run ends. It does not include the LLM turn of the agent, so it is not the total cost of the run. Amounts are in USD micros, so divide by 1,000,000 for dollars.

FieldReads asUse
usage.cap.limit_usd_microsThe effective cap of this runCompare with the cap you asked for
usage.cap.counted_usd_microsWhat counts against the cap nowSame value as billable_amount_usd_micros
usage.cap.remaining_usd_microsRoom left under the capAlert when it gets small
usage.debited_usd_microsWhat the wallet really deducted for the run and threadCost accounting. Includes the LLM row

The effective cap comes from the request. With nothing sent it is the Format's generation_spend_cap_usd_micros, or $400 when the Format never named one. A number up to 500 is accepted as written, null means the $500 platform maximum, and 0 or a value above 500 returns 400.

A headroom check

The function below reads the three fields from a receipt's usage and returns the percentage used. It returns null when usage is null, which the docs distinguish from 0: null means the API could not read the spend at all. The sample feeds it a made-up receipt fragment; there are no network calls.

// usage.cap on a run receipt: limit, counted (= billable_amount_usd_micros), remaining.
export function capHeadroom(usage) {
  if (!usage?.cap) return null; // usage is null when the API could not read the spend
  const { limit_usd_micros: limit, counted_usd_micros: counted } = usage.cap;
  return { counted, remaining: usage.cap.remaining_usd_micros, usedPct: Math.round((counted / limit) * 100) };
}

const receiptUsage = {
  billable_amount_usd_micros: 90_000_000,
  cap: { limit_usd_micros: 120_000_000, counted_usd_micros: 90_000_000, remaining_usd_micros: 30_000_000 },
};
console.log(capHeadroom(receiptUsage));
console.log(capHeadroom(null));

Using it in a monitor

  • Poll the full receipt, or status_url. A terminal status adds usage to the status payload, and the full receipt carries it at every status.
  • Warn on remaining_usd_micros, not on a percentage you invented. A run that is near its limit during a retry of a single scene may have only a few dollars of room.
  • Do not treat billable_amount_usd_micros as an invoice. The docs call it a receipt value. GET /v1/usage and GET /v1/balance are the billing records.
  • The wallet fields (debited_usd_micros, held_usd_micros, refunded_usd_micros) are null on receipts written before the ledger answered. held and refunded are not spend, and final becomes true when no hold is open.

Cap hit versus wallet empty

Two gates act at different times. The wallet gate is at create: a workspace that cannot fund the run gets 402 insufficient_credits with next_action: add_funds, or 402 organization_wallet_not_provisioned, and nothing ran. The cap gate acts during the run: the run ends failed, and you still pay for generation that finished. A cap failure is a sizing problem, so raise generation_spend_cap_usd for the next run rather than retrying the same value. The Errors and spend page lists the codes, and the Runs and results page lists the usage fields.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume