Watch a Format run's spend against its cap while it runs
A Format run's receipt reports usage.cap with limit, counted and remaining while it is in flight. Read it to see how close a video run is to failing on its cap.

Poll the run receipt and read usage.cap: it carries limit_usd_micros, counted_usd_micros and remaining_usd_micros, and the counted figure climbs while the run is still in flight. When remaining_usd_micros is close to zero, the next paid generation step is the one that can end the run as failed on its cap.
This matters most for Formats that make video. A run is one agent turn that calls generation tools, and a clip from a catalog model such as seedance-2.5 (4 to 30 seconds per the Video Router docs) is a single step that can use a large share of a small cap. Watching the cap is cheaper than finding out from a failure.
What the cap accounting counts
usage.billable_amount_usd_micros is the generation spend attributed to the run. It is the running total the cap is enforced against, and it counts both reserved and captured amounts, so it can rise when a step is reserved, before the step finishes. It excludes the agent's own LLM turn, which means it is not the run's total cost.
usage.cap spells out the same accounting: the limit, the counted amount and what remains. A usage of null means spend could not be read at all, which is different from 0. Treat null as unknown and read again rather than as an empty bill.
| Field | What it tells you |
|---|---|
| usage.cap.limit_usd_micros | The ceiling this run is held to |
| usage.cap.counted_usd_micros | Reserved plus captured generation spend, the figure the cap is checked against |
| usage.cap.remaining_usd_micros | Headroom left before the run would fail on its cap |
| usage.debited_usd_micros | What the wallet actually deducted for the run and its thread, LLM turn included |
| usage.held_usd_micros | Holds still open, not spend yet |
| usage.final | true once no hold is open |
Read it in a poll loop
The sample polls the receipt every 15 seconds and prints the headroom as a percentage. It uses only the standard library and reads a key from the environment. Polling spends the read budget, which is separate from and larger than the write budget, so a 15 second interval is well inside any plan.
Stop on a terminal status. completed and failed both end the loop; a run that is failed after spending to its cap lands on the generic format_run_failed code, so compare counted against the limit before you raise the brief.
import json, os, time, urllib.request
BASE = "https://api.sume.com/v1/format-runs/"
KEY = os.environ["SUME_API_KEY"]
def receipt(run_id):
req = urllib.request.Request(BASE + run_id, headers={"Authorization": "Bearer " + KEY})
with urllib.request.urlopen(req) as resp:
return json.load(resp)["data"]
def watch(run_id):
while True:
run = receipt(run_id)
cap = (run.get("usage") or {}).get("cap")
if cap and cap["limit_usd_micros"]:
left = cap["remaining_usd_micros"] / cap["limit_usd_micros"]
print(run["status"], f"{left:.0%} of the cap left")
else:
print(run["status"], "usage unavailable, read again")
if run["status"] not in ("queued", "processing"):
return run
time.sleep(15)
Choose the cap before the run, not during
The docs describe the cap as set up front: by generation_spend_cap_usd on the request, or inherited from the Format, whose cap defaults to $400 when it never set one. A number above the Format's cap is honored, null runs at the $500 platform maximum, and 0 is rejected. A single-scene retry needs a few dollars; a long production run is typically created with a cap around $120.
If the watcher shows a run burning through its headroom, you have two real options: cancel it with POST /v1/format-runs/{run_id}/cancel, which bills what already finished, or let it fail and continue it with previous_run_id and a fresh cap so the finished clips are not regenerated.
Where the receipt and the ledger meet
The receipt figure is for steering a run, not for invoicing. The cost of a run is usage.debited_usd_micros, and the same rows sit behind GET /v1/usage?run_id=, so the receipt, the ledger and an answer from an agent never disagree. Use the cap fields to decide what to do next, and the debited figure to book the cost once usage.final is true.
Sources
Related posts
More in Formats
- Format run webhook retries: dedupe on request_id, order by created_at
Sume Format run webhook retries repeat request_id, which equals run_id. Dedupe on it and order deliveries by created_at, which changes per built body.
- Fruits Drama Format: talking produce for a Thanksgiving push
Call Sume's fruits-drama Format for about 5 second vertical clips where a fruit character speaks your line, queued as five runs for the produce aisle.
- Green-screen Format reaction clip, revised with previous_run_id
Call Sume's green-screen Format for a talking-head overlay ad, then fix one thing with previous_run_id instead of paying for a full re-run.
- Magazine cover and model portrait Formats for beauty sellers
Two Sume Formats make 9:16 beauty images from a packshot: a magazine-cover campaign and a model product portrait. When each helps and what to verify.
Written by Sume