Format run cap headroom: usage.cap limit, counted and remaining
usage.cap on a run receipt splits the spend-cap check into limit, counted and remaining USD micros. Read it to see how close a run is to failing.

To see how near a Format run is to its spend cap, read usage.cap on the receipt. It gives the cap check in three parts: limit_usd_micros, counted_usd_micros and remaining_usd_micros. counted_usd_micros is the same number as usage.billable_amount_usd_micros. When the counted amount reaches the limit, the run stops spending and ends as failed, and usage shows how near the cap it got.
What each number is
Sume enforces the cap of a run against billable_amount_usd_micros. That total counts reserved and captured generation amounts, rises while the run is in progress and settles when the run ends. It does not include the LLM turn of the agent, so it is not the total cost of the run. Amounts are in USD micros, so divide by 1,000,000 for dollars.
| Field | Reads as | Use |
|---|---|---|
usage.cap.limit_usd_micros | The effective cap of this run | Compare with the cap you asked for |
usage.cap.counted_usd_micros | What counts against the cap now | Same value as billable_amount_usd_micros |
usage.cap.remaining_usd_micros | Room left under the cap | Alert when it gets small |
usage.debited_usd_micros | What the wallet really deducted for the run and thread | Cost accounting. Includes the LLM row |
The effective cap comes from the request. With nothing sent it is the Format's generation_spend_cap_usd_micros, or $400 when the Format never named one. A number up to 500 is accepted as written, null means the $500 platform maximum, and 0 or a value above 500 returns 400.
A headroom check
The function below reads the three fields from a receipt's usage and returns the percentage used. It returns null when usage is null, which the docs distinguish from 0: null means the API could not read the spend at all. The sample feeds it a made-up receipt fragment; there are no network calls.
// usage.cap on a run receipt: limit, counted (= billable_amount_usd_micros), remaining.
export function capHeadroom(usage) {
if (!usage?.cap) return null; // usage is null when the API could not read the spend
const { limit_usd_micros: limit, counted_usd_micros: counted } = usage.cap;
return { counted, remaining: usage.cap.remaining_usd_micros, usedPct: Math.round((counted / limit) * 100) };
}
const receiptUsage = {
billable_amount_usd_micros: 90_000_000,
cap: { limit_usd_micros: 120_000_000, counted_usd_micros: 90_000_000, remaining_usd_micros: 30_000_000 },
};
console.log(capHeadroom(receiptUsage));
console.log(capHeadroom(null));Using it in a monitor
- Poll the full receipt, or
status_url. A terminal status addsusageto the status payload, and the full receipt carries it at every status. - Warn on
remaining_usd_micros, not on a percentage you invented. A run that is near its limit during a retry of a single scene may have only a few dollars of room. - Do not treat
billable_amount_usd_microsas an invoice. The docs call it a receipt value.GET /v1/usageandGET /v1/balanceare the billing records. - The wallet fields (
debited_usd_micros,held_usd_micros,refunded_usd_micros) arenullon receipts written before the ledger answered.heldandrefundedare not spend, andfinalbecomestruewhen no hold is open.
Cap hit versus wallet empty
Two gates act at different times. The wallet gate is at create: a workspace that cannot fund the run gets 402 insufficient_credits with next_action: add_funds, or 402 organization_wallet_not_provisioned, and nothing ran. The cap gate acts during the run: the run ends failed, and you still pay for generation that finished. A cap failure is a sizing problem, so raise generation_spend_cap_usd for the next run rather than retrying the same value. The Errors and spend page lists the codes, and the Runs and results page lists the usage fields.
Sources
Related posts
More in Developers
- Validate Wan 3.0 reference limits in Python before submitting
wan-3.0 takes 10 images, 5 videos (15 s total, 16 fps minimum) and 5 audios (15 s total). A short Python check catches an over-limit manifest early.
- verifyWebhook in a fetch handler: four rules, 204 for unknown events
Use @sume-com/sdk verifyWebhook on the raw body, await it, treat false as 401 and answer unknown events with 204. A runnable handler for Workers, Deno and Node.
- Wan 3.0 API request cheat sheet: three modes, 2 to 30 seconds
Wan 3.0 on Sume: the request body for text, first/last frame and reference modes, the 480p/720p/1080p rates and the 2 to 30 second window, on one page.
- GPT Image 1 to GPT Image 2.5 on Sume: what changes in the output
Moving from GPT Image 1 to ChatGPT Image 2.5 on Sume changes the response (URL, not base64), default quality, size grid and failures.
Written by Sume