Format run spend cap: above the Format cap is honored, null is $500

generation_spend_cap_usd on a Sume Format run may exceed the Format's cap and is not clamped; null runs at the $500 maximum; 0 or over 500 is a 400.

5 min readSume
All posts

On a Sume Format run, generation_spend_cap_usd sets that run's own ceiling. A number up to 500 is used as sent, even when it is higher than the Format's own cap, because the docs say it is honored, not clamped. null runs at the $500 platform maximum, which lifts the ceiling without removing it, and 0 or anything above 500 is a 400.

All of it is from the Spend caps section of Create a run, echoed on Errors and spend. The effective cap comes back on every receipt, so you can check what a run actually got.

What cap does my run get?

Each case from the docs:

Effective cap for a Format run, from Sume's Create a run page (read 2026-10-03)
You sendThe run's cap
NothingThe Format's own cap (the platform default is $400 when the Format never named one)
A number up to 500That number, even above the Format's cap
nullThe platform maximum, $500
0, or above 500400: a run that cannot spend cannot deliver

Where do I read the cap and the spend?

The Format's cap is generation_spend_cap_usd_micros on GET /v1/formats/.... On the receipt, usage.generation_spend_cap_usd_micros is the effective cap for the run, and usage.billable_amount_usd_micros is what counts against it so far. Both are in USD micros, so 120000000 is $120.

The billable figure counts reserved and captured amounts, climbs while the run is in flight, and excludes the agent's own LLM turn. It is a receipt figure, not an invoice; GET /v1/usage and GET /v1/balance are the billing records.

What happens when a run hits its cap?

The run ends failed. The errors page says a run that wanted to spend past its cap lands on the generic format_run_failed, so compare usage.billable_amount_usd_micros with the cap before you blame the brief. Generation that finished before the stop is billed, and artifacts[] still lists what was made.

To finish the job, continue the failed run with previous_run_id and a fresh cap rather than starting over; finished clips are already on the thread and are not regenerated.

How should I choose a number?

The docs say production live-commerce integrations run with caps around $120, and a single-scene retry on the same thread needs a fraction of that. Set the cap from the thing being made, not from a blanket maximum.

Reserve null for cases where you have decided that $500 per run is acceptable. A cap is the one spend control you own on the request, and it is enforced during the run, while the wallet check at create only decides whether the run may start (402 insufficient_credits if not).

Sources

Related posts

More in Formats

All Formats posts

Written by Sume