Format run spend cap: $400 default, $500 ceiling, and what is debited

A Format run's generation_spend_cap_usd defaults to the Format's cap ($400 if unset), maxes at $500, rejects 0. billable_amount meters the cap; debited is cost.

5 min readSume
All posts

A Format run can never spend more than its effective generation cap: the Format's own cap, which is $400 when the Format never set one, or the value you send as generation_spend_cap_usd on the run, up to the platform maximum of $500. null means $500, and 0 or anything above 500 is rejected with a 400. What the run cost is not the cap and not the receipt's billable amount; it is usage.debited_usd_micros.

This distinction trips up budget reports, so the rest of this post lays out the three numbers side by side.

What you send and what the cap becomes

The call docs give the mapping.

generation_spend_cap_usd on a Format run, docs read 2026-10-09
You sendThe run's cap
NothingThe Format's cap ($400 if the Format named none)
A number up to 500That number, even above the Format's cap
null$500, the platform maximum
0, or above 500400 error

Three numbers on the receipt

usage.billable_amount_usd_micros is the generation spend counted against the cap: reserved plus captured generation rows, without the agent's LLM turn. It rises while the run is in progress and settles when the run ends. It is a meter for the cap, not the cost.

usage.debited_usd_micros is the real amount the wallet deducted for the run and its thread, including captured rows of every type and the LLM row of the turn. held_usd_micros are open holds, and refunded_usd_micros are holds given back. final is true when no hold is open. GET /v1/usage?run_id= returns the same fold, so the receipt, the ledger and an agent's answer agree.

usage itself is null when the API could not read the spend. That is not the same as zero.

A worked example

Say a run has a $120 cap, spends $85 on generation that captured, and has $20 held for a job still processing. The cap meter reads $105 and the remaining headroom is $15. If the agent's LLM turns cost $1.40, debited_usd_micros is 86,400,000 (the $85 plus the $1.40) while the $20 sits in held_usd_micros. The numbers are illustrative; the field relationships come from the docs.

If the held job then fails, its hold is refunded: the meter drops by $20 and refunded_usd_micros rises by 20,000,000, and the run frees headroom for a retry.

Failure and what you still pay

You pay for generation that finished before a cancel or a failure. If a later step fails you do not get a refund for that earlier generation. A 4xx at create, an idempotent 200 replay and a skipped run cost nothing. Production live-commerce integrations run with caps around $120 in the docs, and a single-scene retry needs a fraction of that, so set the cap per run rather than relying on the Format default of $400.

Choosing a cap

Pick the cap from the most expensive plausible run, not the average. A single-scene retry needs a few dollars; a long-form production run sits near $120 in the docs. A cap of $150 on a $120 workflow leaves room for one retry and stops a runaway loop at $150. Sume accepts a request value above the Format's own cap and does not clamp it, so a client that sets 500 on every call overrides a conservative Format default. Keep the cap in the client and log usage.cap from each receipt.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume