usage.cap on a Format run: an agent reads remaining_usd_micros
A Format run receipt splits its spend cap into limit, counted and remaining USD micros. A supervising agent can read headroom before it asks for more work.

A supervising agent can read how much of a run's cap is left from the Format run receipt. The docs say usage.cap gives the cap calculation in parts: limit_usd_micros, counted_usd_micros and remaining_usd_micros. Counted is the same value as billable_amount_usd_micros, which Sume enforces the cap against and which counts reserved and captured amounts (Format runs).
Units and a worked example
Amounts are in USD micros, where 1,000,000 micros is $1. Take a run with a cap of $8, which is 8,000,000 micros. If counted is 5,250,000, then remaining is 8,000,000 - 5,250,000 = 2,750,000, or $2.75. Those figures are an example I made up for the arithmetic, not a real run.
The counted value rises while the run is in flight and settles when the run ends. So a mid-run reading is a snapshot, and it can only grow.
| Field | Meaning | Spend or not |
|---|---|---|
| billable_amount_usd_micros | Generation spend the cap is enforced against | Reserved plus captured |
| usage.cap.limit_usd_micros | The cap for the run | Limit |
| usage.cap.counted_usd_micros | Same as billable amount | Counted |
| usage.cap.remaining_usd_micros | Limit minus counted | Headroom |
| debited_usd_micros | What the wallet actually deducted, including the LLM row | Real cost |
Counted is not the cost
The docs are firm that the cap's counted value does not include the agent's own LLM turn, so it is not the total cost of the run. The cost is debited_usd_micros, the amount the wallet deducted for the run and its thread. held_usd_micros are open holds, not spend yet, and final turns true when no hold is open.
usage is null when the API could not read the spend at all. That is different from zero, so do not treat a null as free.
A headroom check
The function below decides whether to ask for another step. It treats a null usage as unknown and refuses. It runs offline.
def headroom_usd(receipt: dict):
usage = receipt.get("usage")
if not usage or not usage.get("cap"):
return None
return usage["cap"]["remaining_usd_micros"] / 1_000_000
def can_ask(receipt: dict, next_step_usd: float) -> bool:
left = headroom_usd(receipt)
return left is not None and left >= next_step_usd
if __name__ == "__main__":
r = {"usage": {"cap": {"limit_usd_micros": 8_000_000,
"counted_usd_micros": 5_250_000,
"remaining_usd_micros": 2_750_000}}}
print(headroom_usd(r), can_ask(r, 3.0), can_ask({"usage": None}, 1))Reading the receipt from a run
The same usage object appears wherever the run receipt is returned: in the response of a poll, in the run webhook payload, and in the list of runs. GET /v1/usage?run_id= reads the same rows, so you can reconcile a receipt against the usage view when a number looks off.
A supervising agent should read the receipt fresh before each decision rather than caching it. The counted value only grows during a run, and final stays false until every hold is closed. When final is true and held_usd_micros is zero, the numbers are settled and safe to book.
Checklist
Use the receipt rather than your own running total.
- Read
remaining_usd_micros, do not compute it yourself. - Treat null
usageas unknown. - Reconcile cost with
debited_usd_microsorGET /v1/usage. - You can lower a cap on a later run, but one run cannot raise its own.
Sources
Related posts
More in Agents
- Sume hosted MCP is not Studio Agent: the one URL to put in your client
Put https://mcp.sume.com/mcp into Claude, Cursor or Codex. Studio Agent is the in-app product, not a customer MCP server, so there is no second URL to add.
- Sume jobs_wait limits: 20 ids and 55 seconds per call
Hosted Sume MCP jobs_wait takes at most 20 job ids of 256 characters and waits up to 55 seconds. How to wait on a 30-clip storyboard in groups.
- Launch-day run of a Sume schedule: one key per model, input as data
Start a schedule run the day a model launches with POST /v1/actions/{id}/runs. Send the model name as input data and derive the Idempotency-Key from it.
- Make a video from a URL in Claude or Cursor: crawl_scrape, Wan 3.0
Alibaba's Wan 3.0 reads webpages; Sume's request has no web_url. In Claude or Cursor use hosted MCP: crawl_scrape the page, then generate_video with wan-3.0.
Written by Sume