Why Sume usage rows add up to more than you spent
Summing rows from GET /v1/usage overstates spend: refunded rows keep their hold amount. Read summary.debited_usd for a thread, run or job instead.

If you add up the amounts on the rows from Sume's GET /v1/usage you will usually get a number higher than what left your wallet, because a refunded row keeps the amount of its hold. The figure to quote is summary.debited_usd on a scoped usage read, and the Usage docs say never to sum rows yourself.
This post explains the three statuses, what the summary folds for you, and a short script that prints spend for one run.
What do reserved, captured and refunded mean?
Sume reserves the estimated amount when a paid job is accepted, captures it when the job completes, and releases it if the job fails or is cancelled before capture. The usage ledger records each state as a status.
A single job that fails therefore leaves a row whose amount equals the original hold, marked refunded. Adding that row to the total counts money that was never taken.
| Status | Meaning | Is it spend? |
|---|---|---|
reserved | Estimated usage held before provider execution | No, still a hold |
captured | Billable usage taken after successful completion | Yes |
refunded | Reserved usage released after failure or cancellation | No |
What does the scoped summary add?
Add thread_id, run_id or job_id to the usage call and the response gains a summary folded over every ledger row the scope caused. The limit parameter only caps the rows listed, not the summary.
Read final first. If it is false, a hold is open and debited_usd can still rise.
| Field | What it means |
|---|---|
debited_usd_micros, debited_usd | What the wallet deducted: captured rows of every operation type. The figure to quote. |
held_usd_micros | Holds still open. Not spend yet. |
refunded_usd_micros | Holds given back after a failure, cancellation or queue_full. Not spend. |
final | true once no hold is open. |
How do I read one run's cost?
The script below calls the usage endpoint with a run id and prints the three money fields. It uses only the standard library and reads the key and run id from the environment.
import json, os, urllib.request
def usage_summary(run_id: str) -> dict:
req = urllib.request.Request(
f"https://api.sume.com/v1/usage?run_id={run_id}&limit=50",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
)
with urllib.request.urlopen(req) as resp:
return json.load(resp)["summary"]
summary = usage_summary(os.environ["RUN_ID"])
print("spent:", summary["debited_usd"])
print("still held:", summary["held_usd_micros"] / 1_000_000)
print("refunded:", summary["refunded_usd_micros"] / 1_000_000)
print("final:", summary["final"])What does a scope include besides generation jobs?
For a Studio Agent thread or an agent run, the docs say the summary covers every row the scope caused, and the includes object counts rows by kind: LLM turns, sidecars, Browser sessions, generation jobs, script_run_children (jobs that a script_run call dispatched), and refunded rows. by_operation_type splits the same money by operation type.
This is why a thread's cost is more than the sum of its video jobs. The agent's own turns are in debited_usd too.
Is this the same as the run's spend cap?
No. For a run, the response also carries a cap block: what counts against the run's generation cap, which is reserved plus captured generation rows with LLM excluded, and what is left. The docs say it is never a cost. Use debited_usd for what you spent and cap for how much room remains.
For plan changes and top-ups, use Billing and credits.
What should I put in a monthly report?
Use scoped reads for anything you attribute to a project, and the balance for anything you reconcile against a payment. A report that sums raw rows will drift from the wallet by the amount of every refund.
If you must work from rows, filter to captured ones and then check the total against debited_usd for a sample scope. Where they disagree, trust the summary and look for the rows you misclassified.
For the whole workspace, read the balance before and after a period and compare it with the top-ups and grants in the ledger.
Sources
Related posts
More in Pricing
- xAI Batch API discount: 20% on four text models, none on images
xAI's pricing page says the Batch API discount applies to text models only, 20% on four listed ones. Image and video batches are billed at standard rates.
- How Sume pricing works: plans, one wallet, published model rates
Sume plans set access and concurrency. Usage draws from one prepaid wallet at each model's published USD rate, for generation, the Agent, Formats, and the API.
- AI avatar video API pricing: cost per second and per minute
Sume bills AI avatar video per second by quality tier, with separate rates when you send a product image. Per-minute costs for standard, plus, and max.
- Estimate AI video generation cost before running a Sume job
See what an AI video will cost on Sume before paying: published rates, GET /v1/catalog estimates, unbilled plan checks, MCP dry runs, and spend caps.
Written by Sume