generation_spend_cap_usd: null means $500, 0 is a 400, per ad variant

What the per-run spend cap does on Sume Format runs for a number, null, 0 and a value over 500, and why every ad variant should set its own.

5 min readSume
All posts

On a Sume Format run, generation_spend_cap_usd is the ceiling for that one run. A number up to 500 is accepted as written, even if it is above the Format's own cap. null means the platform maximum of $500. 0 and anything above 500 return 400. If you send nothing, the run inherits the Format's cap, and a Format that never set one reports the platform default of $400. For an ad batch, set the cap yourself on every item, because the inherited number is likely much larger than one ad needs.

The outcomes

The table is from the Create a run page. Read it once and write it into your client as a validation step.

Effective cap by what you send (Sume docs, read 2026-10-05)
You sendThe run's cap
NothingThe Format's cap, or $400 if it never named one
A number up to 500That number, even above the Format's cap
null$500, the platform maximum
0400 invalid_request
A number above 500400 invalid_request

What happens when the cap is hit

The cap is checked during the run. A run cannot spend more than its effective cap, and a run that reaches it ends as failed, with usage showing how near the spend got. Media that was already made stays in artifacts[] on a failed run. You pay for generation that finished before a failure, and a later failure does not refund it. A 4xx at create, an idempotent replay and a skipped run cost nothing.

A second gate sits at create. The workspace wallet must be able to fund the run, or the create fails with 402 insufficient_credits and nothing runs. The cap protects you from a long run. The wallet gate protects you from an empty account. They fail at different times.

Choose a number for an ad

The docs give two anchors from production: long-form runs usually carry caps of approximately $120, and a single-scene retry carries a few dollars. Do not borrow those numbers for a short ad. Run one ad with a generous cap, read usage.billable_amount_usd_micros on the receipt, and set the cap for the batch at a multiple of that number that you are willing to lose on a bad item.

In a bulk queue each item carries its own cap, so build the items with one.

import json

CAP_USD = 25
items = [
    {"instruction": f"15 s vertical ad, hook {n}", "generation_spend_cap_usd": CAP_USD}
    for n in range(1, 6)
]
body = {"concurrency": 4, "items": items}
assert all(0 < i["generation_spend_cap_usd"] <= 500 for i in items)
print(json.dumps(body)[:120])

Reading the number back

The receipt's usage.billable_amount_usd_micros is the spend of the run against the cap. It rises while the run is in progress, counts reserved and captured amounts, and settles when the run ends. It does not include the agent's own LLM turn, so it is not the total cost of the run, and it is a receipt value and not an invoice. GET /v1/usage and GET /v1/balance are the billing records. usage is null when the API could not read the spend, which is different from 0. Log it for every run, and compare it with the cap you set before you raise the cap.

Do not use null as a way to stop thinking about caps. It lifts the ceiling to $500 and does not remove it, so a stuck run can still spend up to that figure on one ad. Use null only for a run that you have watched and that really needs the room, and write down why. For routine variants, a small explicit number is better, because a run that ends as failed at the cap is a cheap signal, while a run that quietly spends $300 on a 15-second ad is not.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume