How to cap a Sume Format run's spend: $500 ceiling, $400 default
Send generation_spend_cap_usd on each Format run. A value above 500 returns 400, a Format with no cap defaults to 400, and a run past its cap fails.

Set generation_spend_cap_usd on the Format run request, in dollars. The platform maximum is $500, and a value above $500 gets a 400. If a Format never named a cap, its default is $400. A run that tries to spend more than its cap does not run over. It ends as failed with the generic code format_run_failed, so a cap is a real stop, not a warning.
Pick the cap before you send the request, not after a surprise on the usage page. The cap is the one number that bounds the worst case of a single call, and it is also the number a finance reviewer will ask about first.
What the cap does and does not do
A cap is per run, not per day and not per account. Sume accepts a number above the Format's own cap, up to the platform ceiling, so the request can raise or lower the limit that the Format author set. That makes the field a tool for the caller. If you embed a Format in your own product, pick the cap from what the customer paid for, and not from a single number in your config.
There is one more reason to set it explicitly. A Format that never named a cap falls back to $400, and that is far above what many single runs need. Leaving the default in place means your real exposure per run is $400, even when a typical run costs much less.
Where to read the cap
Read the effective cap from two places. PublicFormat.generation_spend_cap_usd_micros shows the Format level value and is always a number. The run receipt shows usage.generation_spend_cap_usd_micros, which is the cap that applied to that run. Micros are dollars times 1,000,000, so $400 is 400,000,000 micros.
| Field | Where | Meaning |
|---|---|---|
| generation_spend_cap_usd | Run request | Per run ceiling in dollars, max 500 |
| generation_spend_cap_usd_micros | PublicFormat | Format level cap, default 400 dollars |
| usage.generation_spend_cap_usd_micros | Run receipt | The cap that applied to this run |
| usage.billable_amount_usd_micros | Run receipt | What the run billed, in micros |
Validate before you send
Validate the number before you send it. The helper below mirrors the documented rules and converts a dollar cap into micros for comparison with a receipt. It uses Decimal, so there is no float rounding.
from decimal import Decimal
PLATFORM_MAX = Decimal("500")
FORMAT_DEFAULT = Decimal("400")
def effective_cap(requested=None):
cap = FORMAT_DEFAULT if requested is None else Decimal(str(requested))
if cap <= 0 or cap > PLATFORM_MAX:
raise ValueError("generation_spend_cap_usd must be above 0 and at most 500")
return cap
def to_micros(usd):
return int(usd * 1_000_000)
for ask in (None, 3, 500):
cap = effective_cap(ask)
print(ask, cap, to_micros(cap))
Diagnose a run that keeps failing
If runs of one plan tier fail again and again, compare the receipt cap first. The docs call this out because a run that tries to spend more than its cap fails with the same code as other generic failures. Compare usage.billable_amount_usd_micros and usage.generation_spend_cap_usd_micros on the failed receipts, and raise the request cap only if the plan allows it.
The ledger helps afterwards. GET /v1/usage lists rows with the statuses reserved, captured and refunded, and it can be filtered by run_id or job_id. Its summary gives debited_usd_micros, so you can compare what a run actually cost with the cap you sent.
Combine the cap with an idempotency key
Use a cap together with an idempotency key. A retry with the same key and body returns the original receipt with idempotency_hit: true, and does not start a second paid run. A retry of a failed run needs a new key, and you can send a higher cap with it. If you change only the cap on a retry that keeps the old key, the body differs, and you get a 409 idempotency_conflict.
Keep one more rule in mind. A cap is not a price quote. It says the most Sume may spend on the run, and the receipt says what the run billed. Log both next to your own order id, so a support reply can state the cap, the billed amount and the outcome in one line.
Sources
Related posts
More in Developers
- Retry a lip-sync submit after a timeout without double billing
A timed-out submit may or may not have created a job. Resend the same body with the same Idempotency-Key and Sume returns the same job, not a second charge.
- Idempotency-Key per shot: rerun one failed AI video shot in Python
One key per shot, derived from project, shot number and revision: a retry returns the same Sume job, a changed prompt gets a new key. Python key helper inside.
- Chain Ideogram 4.5 edits with webhooks: job.completed starts pass 2
Run a multi-turn Ideogram 4.5 edit chain on Sume without polling: submit with mode webhook, verify the signature, and start the next pass from job.completed.
- Ideogram 4.5 seed on Sume returns 400: how to repeat an edit
Ideogram's own API takes a seed for 4.5 edits, but Sume returns 400 unsupported_parameter for seed on every image model. Keep the output URL, not the seed.
Written by Sume