Format run spend cap: the $400 default, in receipt micros

A Format with no cap set reports 400000000 micros. How the run cap, the $500 maximum, null and 0 behave, and why billable_amount is not the whole bill.

5 min readSume
All posts

If a Format never set a spend cap, the platform default is $400, which appears as 400000000 in generation_spend_cap_usd_micros. One micro is one millionth of a dollar. A run can never spend more than its own effective cap, and you can set that cap per run with generation_spend_cap_usd, up to $500.

How the cap resolves

The effective cap is decided at create. The table lists every case the call docs describe, as of 2026-10-09.

Effective run cap by what you send, as of 2026-10-09
You sendEffective capIn micros
NothingThe Format's cap (default $400)400000000 if the Format set none
generation_spend_cap_usd: 25$2525000000
A number above the Format's cap, up to 500That number; not clampede.g. 450000000
nullThe $500 platform maximum500000000
0, or above 500400 invalid_requestnone

Reading the receipt

Every receipt gives the cap as usage.generation_spend_cap_usd_micros and the counted spend as usage.billable_amount_usd_micros. usage.cap splits the same calculation into limit_usd_micros, counted_usd_micros and remaining_usd_micros. For a $25 cap with 4,250,000 micros counted, the remaining value is 25,000,000 - 4,250,000 = 20,750,000 micros, or $20.75.

The counted value includes reserved and captured amounts, and it rises while the run is in progress. It does not include the agent's own LLM turn. The amount the wallet really paid is usage.debited_usd_micros, so use that for cost reporting and billable_amount_usd_micros for the cap check.

What happens at the cap

A run that tries to spend more than its cap ends failed with the generic format_run_failed code. Before you raise the brief, compare billable_amount_usd_micros with generation_spend_cap_usd_micros. Generation that finished before the failure is still billed, and artifacts[] still lists the media. Create-time errors such as 4xx, an idempotent 200 replay and a skipped run cost nothing.

A practical rule from the docs: long-form production runs often use caps around $120, and a single-scene retry needs a few dollars. Pick a cap close to what you expect to spend, because the cap is the only limit you control per run.

Choosing a number

There is no single right cap. Use the stage of the work to choose one. For a first test of a new Format, a low cap such as $5 to $10 shows quickly whether the brief is sound and limits the cost of a mistake. For production, read usage.billable_amount_usd_micros from a few finished runs and set the cap a little above the highest. The docs give production long-form runs at about $120 as an example, which is their figure and not a promise for your Format.

Remember that a cap can be set on the Format and on the run. A caller can send a number above the Format's own cap and it is not clamped, so if you publish a Format that others call, the Format's cap is not a hard guard against a caller who sends a larger number. Use the wallet and the plan's concurrency as the outer limits.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume