Watch a Format run's spend against its cap while it runs

A Format run's receipt reports usage.cap with limit, counted and remaining while it is in flight. Read it to see how close a video run is to failing on its cap.

5 min readSume
All posts

Poll the run receipt and read usage.cap: it carries limit_usd_micros, counted_usd_micros and remaining_usd_micros, and the counted figure climbs while the run is still in flight. When remaining_usd_micros is close to zero, the next paid generation step is the one that can end the run as failed on its cap.

This matters most for Formats that make video. A run is one agent turn that calls generation tools, and a clip from a catalog model such as seedance-2.5 (4 to 30 seconds per the Video Router docs) is a single step that can use a large share of a small cap. Watching the cap is cheaper than finding out from a failure.

What the cap accounting counts

usage.billable_amount_usd_micros is the generation spend attributed to the run. It is the running total the cap is enforced against, and it counts both reserved and captured amounts, so it can rise when a step is reserved, before the step finishes. It excludes the agent's own LLM turn, which means it is not the run's total cost.

usage.cap spells out the same accounting: the limit, the counted amount and what remains. A usage of null means spend could not be read at all, which is different from 0. Treat null as unknown and read again rather than as an empty bill.

Fields on the run receipt usage block (read 2026-10-03)
FieldWhat it tells you
usage.cap.limit_usd_microsThe ceiling this run is held to
usage.cap.counted_usd_microsReserved plus captured generation spend, the figure the cap is checked against
usage.cap.remaining_usd_microsHeadroom left before the run would fail on its cap
usage.debited_usd_microsWhat the wallet actually deducted for the run and its thread, LLM turn included
usage.held_usd_microsHolds still open, not spend yet
usage.finaltrue once no hold is open

Read it in a poll loop

The sample polls the receipt every 15 seconds and prints the headroom as a percentage. It uses only the standard library and reads a key from the environment. Polling spends the read budget, which is separate from and larger than the write budget, so a 15 second interval is well inside any plan.

Stop on a terminal status. completed and failed both end the loop; a run that is failed after spending to its cap lands on the generic format_run_failed code, so compare counted against the limit before you raise the brief.

import json, os, time, urllib.request

BASE = "https://api.sume.com/v1/format-runs/"
KEY = os.environ["SUME_API_KEY"]


def receipt(run_id):
    req = urllib.request.Request(BASE + run_id, headers={"Authorization": "Bearer " + KEY})
    with urllib.request.urlopen(req) as resp:
        return json.load(resp)["data"]


def watch(run_id):
    while True:
        run = receipt(run_id)
        cap = (run.get("usage") or {}).get("cap")
        if cap and cap["limit_usd_micros"]:
            left = cap["remaining_usd_micros"] / cap["limit_usd_micros"]
            print(run["status"], f"{left:.0%} of the cap left")
        else:
            print(run["status"], "usage unavailable, read again")
        if run["status"] not in ("queued", "processing"):
            return run
        time.sleep(15)

Choose the cap before the run, not during

The docs describe the cap as set up front: by generation_spend_cap_usd on the request, or inherited from the Format, whose cap defaults to $400 when it never set one. A number above the Format's cap is honored, null runs at the $500 platform maximum, and 0 is rejected. A single-scene retry needs a few dollars; a long production run is typically created with a cap around $120.

If the watcher shows a run burning through its headroom, you have two real options: cancel it with POST /v1/format-runs/{run_id}/cancel, which bills what already finished, or let it fail and continue it with previous_run_id and a fresh cap so the finished clips are not regenerated.

Where the receipt and the ledger meet

The receipt figure is for steering a run, not for invoicing. The cost of a run is usage.debited_usd_micros, and the same rows sit behind GET /v1/usage?run_id=, so the receipt, the ledger and an answer from an agent never disagree. Use the cap fields to decide what to do next, and the debited figure to book the cost once usage.final is true.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume