Format run spend cap: the $400 default, in receipt micros
A Format with no cap set reports 400000000 micros. How the run cap, the $500 maximum, null and 0 behave, and why billable_amount is not the whole bill.

If a Format never set a spend cap, the platform default is $400, which appears as 400000000 in generation_spend_cap_usd_micros. One micro is one millionth of a dollar. A run can never spend more than its own effective cap, and you can set that cap per run with generation_spend_cap_usd, up to $500.
How the cap resolves
The effective cap is decided at create. The table lists every case the call docs describe, as of 2026-10-09.
| You send | Effective cap | In micros |
|---|---|---|
| Nothing | The Format's cap (default $400) | 400000000 if the Format set none |
generation_spend_cap_usd: 25 | $25 | 25000000 |
| A number above the Format's cap, up to 500 | That number; not clamped | e.g. 450000000 |
null | The $500 platform maximum | 500000000 |
0, or above 500 | 400 invalid_request | none |
Reading the receipt
Every receipt gives the cap as usage.generation_spend_cap_usd_micros and the counted spend as usage.billable_amount_usd_micros. usage.cap splits the same calculation into limit_usd_micros, counted_usd_micros and remaining_usd_micros. For a $25 cap with 4,250,000 micros counted, the remaining value is 25,000,000 - 4,250,000 = 20,750,000 micros, or $20.75.
The counted value includes reserved and captured amounts, and it rises while the run is in progress. It does not include the agent's own LLM turn. The amount the wallet really paid is usage.debited_usd_micros, so use that for cost reporting and billable_amount_usd_micros for the cap check.
What happens at the cap
A run that tries to spend more than its cap ends failed with the generic format_run_failed code. Before you raise the brief, compare billable_amount_usd_micros with generation_spend_cap_usd_micros. Generation that finished before the failure is still billed, and artifacts[] still lists the media. Create-time errors such as 4xx, an idempotent 200 replay and a skipped run cost nothing.
A practical rule from the docs: long-form production runs often use caps around $120, and a single-scene retry needs a few dollars. Pick a cap close to what you expect to spend, because the cap is the only limit you control per run.
Choosing a number
There is no single right cap. Use the stage of the work to choose one. For a first test of a new Format, a low cap such as $5 to $10 shows quickly whether the brief is sound and limits the cost of a mistake. For production, read usage.billable_amount_usd_micros from a few finished runs and set the cap a little above the highest. The docs give production long-form runs at about $120 as an example, which is their figure and not a promise for your Format.
Remember that a cap can be set on the Format and on the run. A caller can send a number above the Format's own cap and it is not clamped, so if you publish a Format that others call, the Format's cap is not a hard guard against a caller who sends a larger number. Use the wallet and the plan's concurrency as the outer limits.
Sources
Related posts
More in Formats
- How long to poll a Format run: use expires_at, not your own timeout
A non-terminal Format run receipt carries expires_at, 90 minutes after created_at or earlier if the run goes silent. Set your poll ceiling from it and back off.
- Format run media URLs never expire and anyone can open them
Media from a Sume Format run lives at durable media.sume.com URLs that anyone with the link can open. What that means for per-customer access and retention.
- Format run spend caps: $400 default, $500 maximum, and what null does
How Sume Format run spend caps work: the $400 platform default, a per-run generation_spend_cap_usd up to $500, null for $500, and 0 or above 500 as a 400.
- Format webhook: 10 s timeout, 10 attempts, about 3 h 5 min of retries
A slow Format webhook gets 10 tries. Without jitter, the last starts 11,010 s after the first; with 10 s timeouts the span is about 3 h 5 min. Code included.
Written by Sume