Format run spend cap: $400 default, $500 ceiling, and what is debited
A Format run's generation_spend_cap_usd defaults to the Format's cap ($400 if unset), maxes at $500, rejects 0. billable_amount meters the cap; debited is cost.

A Format run can never spend more than its effective generation cap: the Format's own cap, which is $400 when the Format never set one, or the value you send as generation_spend_cap_usd on the run, up to the platform maximum of $500. null means $500, and 0 or anything above 500 is rejected with a 400. What the run cost is not the cap and not the receipt's billable amount; it is usage.debited_usd_micros.
This distinction trips up budget reports, so the rest of this post lays out the three numbers side by side.
What you send and what the cap becomes
The call docs give the mapping.
| You send | The run's cap |
|---|---|
| Nothing | The Format's cap ($400 if the Format named none) |
| A number up to 500 | That number, even above the Format's cap |
| null | $500, the platform maximum |
| 0, or above 500 | 400 error |
Three numbers on the receipt
usage.billable_amount_usd_micros is the generation spend counted against the cap: reserved plus captured generation rows, without the agent's LLM turn. It rises while the run is in progress and settles when the run ends. It is a meter for the cap, not the cost.
usage.debited_usd_micros is the real amount the wallet deducted for the run and its thread, including captured rows of every type and the LLM row of the turn. held_usd_micros are open holds, and refunded_usd_micros are holds given back. final is true when no hold is open. GET /v1/usage?run_id= returns the same fold, so the receipt, the ledger and an agent's answer agree.
usage itself is null when the API could not read the spend. That is not the same as zero.
A worked example
Say a run has a $120 cap, spends $85 on generation that captured, and has $20 held for a job still processing. The cap meter reads $105 and the remaining headroom is $15. If the agent's LLM turns cost $1.40, debited_usd_micros is 86,400,000 (the $85 plus the $1.40) while the $20 sits in held_usd_micros. The numbers are illustrative; the field relationships come from the docs.
If the held job then fails, its hold is refunded: the meter drops by $20 and refunded_usd_micros rises by 20,000,000, and the run frees headroom for a retry.
Failure and what you still pay
You pay for generation that finished before a cancel or a failure. If a later step fails you do not get a refund for that earlier generation. A 4xx at create, an idempotent 200 replay and a skipped run cost nothing. Production live-commerce integrations run with caps around $120 in the docs, and a single-scene retry needs a fraction of that, so set the cap per run rather than relying on the Format default of $400.
Choosing a cap
Pick the cap from the most expensive plausible run, not the average. A single-scene retry needs a few dollars; a long-form production run sits near $120 in the docs. A cap of $150 on a $120 workflow leaves room for one retry and stops a runaway loop at $150. Sume accepts a request value above the Format's own cap and does not clamp it, so a client that sets 500 on every call overrides a conservative Format default. Keep the cap in the client and log usage.cap from each receipt.
Sources
Related posts
More in Pricing
- 40 hero images from five premium Sume rows: $7.50 to $15.00
40 hero images cost $7.50 on Nano Banana Pro 1K, $8.00 on 2.1 at 4K, $8.90 on GPT Image 2.5 high 4K, $11.00 on Ideogram 4.5 high and $15.00 on Pro 4K.
- Vertical hook at the Seedance 2.5 floor: 4 s costs $5.69 at 1080p
The shortest Seedance 2.5 job, 4 seconds at 9:16, costs $1.07 at 480p, $2.31 at 720p and $5.69 at 1080p on Sume.
- Four-week video budget: Omni hooks, H3 b-roll, two Recast swaps
A four-week plan with 3 Omni hooks and 2 H3 b-roll clips a week plus two 15-second Recast swaps totals $29.25 on Sume. Every line is priced.
- Free plan seventh submit: 429 queue_full is free, 402 can come first
On Sume Free, six paid jobs fit (1 processing, 5 queued). The seventh gets 429 queue_full with its hold released. Which error comes first, and the cost.
Written by Sume