How to read what one agent run cost on Sume with the usage API
Pass a run id to GET /v1/usage and Sume sums every ledger row the run caused. A short Python script prints debited, held and refunded dollars.

GET /v1/usage can return the total spend of one Format, Action or Agent run if you pass the run's id. Per the API types, the response then carries a summary that folds every ledger row the run caused, including its generation jobs and the run thread's own turns, not only the rows listed.
The three numbers to read
debited_usd_micros is the sum of captured rows and is the source-of-truth for what the wallet lost. held_usd_micros is open holds, which are not spend yet. refunded_usd_micros is what was given back. The final flag turns true when no hold is open, so wait for it before you report a cost.
Script
The script reads the key and run id from the environment, fails with a clear message when either is missing, and prints the three amounts in dollars.
import json, os, sys, urllib.request
key = os.environ.get("SUME_API_KEY")
run_id = os.environ.get("RUN_ID")
if not key or not run_id:
sys.exit("Set SUME_API_KEY and RUN_ID first.")
url = f"https://api.sume.com/v1/usage?run_id={run_id}"
req = urllib.request.Request(url, headers={"Authorization": f"Bearer {key}"})
with urllib.request.urlopen(req) as resp:
summary = json.load(resp)["data"]["summary"]
for field in ("debited_usd_micros", "held_usd_micros", "refunded_usd_micros"):
print(f"{field}: ${summary[field] / 1_000_000:.4f}")
print("final:", summary["final"])Why read the summary rather than add rows yourself
The listed rows are capped by limit, but the summary folds every row of the scope, up to 5,000. Adding up the listed rows yourself can miss spend, and it can count a reserved row that never became spend. The summary already separates captured, held and refunded amounts, so it is the safer number to put in a report.
For cost per run in a team, store the summary figures with the run record the moment final becomes true. A later top-up or unrelated job will not change them.
Caveats
- The summary folds up to 5,000 rows; when
truncatedis true the figures cover the newest rows only. - A run still in flight shows holds. Poll until
finalis true. - The same endpoint takes
thread_idandjob_idfor those scopes, and without a scope it returns the newest rows with no summary.
Sources
Related posts
More in Developers
- A Hurl file for a Sume job: submit, poll until terminal, fetch
A plain-text Hurl file chains submit, a retrying status poll on data.terminal, and the result read. Run it in CI with hurl --variable and an Idempotency-Key.
- Idempotency keys for a SKU catalog: SKU, asset, version
A fresh uuid per request makes Sume's Idempotency-Key decorative. Build it from SKU, asset and version so a retried holiday batch is not charged twice.
- Idempotency keys for batch TTS: derive them from what you render
A retry loop that mints random keys pays twice. Hash the sentence, voice, model and settings into the key, and re-running a script returns the first job.
- Ideogram 4.5 edits keep the source shape; GPT Image 2.5 needs auto
On Sume an Ideogram 4.5 edit without aspect_ratio keeps the source image's shape, while other edit calls should send aspect_ratio auto. What to send for each.
Written by Sume