Format run spend cap: above the Format cap is honored, null is $500
generation_spend_cap_usd on a Sume Format run may exceed the Format's cap and is not clamped; null runs at the $500 maximum; 0 or over 500 is a 400.

On a Sume Format run, generation_spend_cap_usd sets that run's own ceiling. A number up to 500 is used as sent, even when it is higher than the Format's own cap, because the docs say it is honored, not clamped. null runs at the $500 platform maximum, which lifts the ceiling without removing it, and 0 or anything above 500 is a 400.
All of it is from the Spend caps section of Create a run, echoed on Errors and spend. The effective cap comes back on every receipt, so you can check what a run actually got.
What cap does my run get?
Each case from the docs:
| You send | The run's cap |
|---|---|
| Nothing | The Format's own cap (the platform default is $400 when the Format never named one) |
| A number up to 500 | That number, even above the Format's cap |
null | The platform maximum, $500 |
0, or above 500 | 400: a run that cannot spend cannot deliver |
Where do I read the cap and the spend?
The Format's cap is generation_spend_cap_usd_micros on GET /v1/formats/.... On the receipt, usage.generation_spend_cap_usd_micros is the effective cap for the run, and usage.billable_amount_usd_micros is what counts against it so far. Both are in USD micros, so 120000000 is $120.
The billable figure counts reserved and captured amounts, climbs while the run is in flight, and excludes the agent's own LLM turn. It is a receipt figure, not an invoice; GET /v1/usage and GET /v1/balance are the billing records.
What happens when a run hits its cap?
The run ends failed. The errors page says a run that wanted to spend past its cap lands on the generic format_run_failed, so compare usage.billable_amount_usd_micros with the cap before you blame the brief. Generation that finished before the stop is billed, and artifacts[] still lists what was made.
To finish the job, continue the failed run with previous_run_id and a fresh cap rather than starting over; finished clips are already on the thread and are not regenerated.
How should I choose a number?
The docs say production live-commerce integrations run with caps around $120, and a single-scene retry on the same thread needs a fraction of that. Set the cap from the thing being made, not from a blanket maximum.
Reserve null for cases where you have decided that $500 per run is acceptable. A cap is the one spend control you own on the request, and it is enforced during the run, while the wallet check at create only decides whether the run may start (402 insufficient_credits if not).
Sources
Related posts
More in Formats
- Format run webhook retries: dedupe on request_id, order by created_at
Sume Format run webhook retries repeat request_id, which equals run_id. Dedupe on it and order deliveries by created_at, which changes per built body.
- Bulk Format runs: 100 items, 16 at once, what completed means
Sume bulk runs take 1 to 100 items at concurrency 1 to 16. A queue marked completed means every item is terminal, not that every item succeeded.
- Format output_schema for partial results: nullable and primary key
How to design a Sume Format output_schema so a run that makes the video but misses a caption still passes: nullable fields, no minItems, and primary_output_key.
- Format run expires_at: 90 minutes from created_at, or sooner
How long a Sume Format run can live, which statuses are terminal, and how to poll the status_url without waiting on an event stream that does not exist.
Written by Sume