Which step cost the most in a Sume thread? Read by_operation_type
GET /v1/usage with thread_id returns a summary with by_operation_type and runs. A Python snippet ranks operation types by spend and shows how to read it.

To see which step cost the most in a Sume Agent thread, call GET /v1/usage?thread_id=... and read summary.by_operation_type. It folds every ledger row the thread caused into spend per operation type, and summary.runs does the same per run. You do not have to page through rows and add them yourself.
This post shows what the summary contains, which field to quote as the cost, and a short Python script that ranks operation types. The field names come from Sume's usage docs; the exact shape of each entry is in the live OpenAPI, so the script reads it defensively.
A thread is the right scope when an agent chose the steps. If you submitted one job yourself, use job_id; if you started a Format, Action or Agent run, use run_id.
What is in the usage summary?
Adding thread_id, run_id or job_id to the usage call adds a summary block computed over every ledger row the scope caused. The limit parameter only caps how many rows are listed, not what the summary sums.
The includes field of the summary adds row counts by kind, such as model turns, sidecars, generation jobs and refunded rows. It is a quick way to see whether a thread was mostly thinking or mostly rendering.
| Field | What it tells you |
|---|---|
| debited_usd_micros, debited_usd | What the wallet deducted. The figure to quote as cost. |
| held_usd_micros | Holds still open. Not spend yet. |
| refunded_usd_micros | Holds given back after failure or cancellation. Not spend. |
| final | True once no hold is open. |
| by_operation_type, runs | The same money split by operation type and, for a thread, by run. |
| cap | For a run, the generation cap and what is left. Never a cost. |
Why not add up the rows myself?
Because a refunded row keeps its hold amount in billable_amount_usd_micros. A naive sum counts money that went back to the wallet. The summary already separates debited, held and refunded money, so quote debited_usd and leave the arithmetic to Sume.
Also note that debited covers every operation type, including the agent's own turns. A thread that looks cheap on generation can still carry model-turn cost, and the breakdown shows it as a separate operation type.
How do I rank the steps in code?
The script below reads the summary for one thread and prints each operation type with its share. It treats by_operation_type as either a mapping or a list, because the schema lives in the OpenAPI and you should check it before relying on the exact shape.
import os
import requests
thread_id = os.environ["SUME_THREAD_ID"]
r = requests.get(
"https://api.sume.com/v1/usage",
params={"thread_id": thread_id, "limit": 1},
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
summary = r.json().get("summary", {})
print("debited:", summary.get("debited_usd"))
print("final:", summary.get("final"))
breakdown = summary.get("by_operation_type")
print(breakdown)What does final: false mean for my report?
It means at least one hold is still open, so the thread is still spending or waiting to settle. Do not publish a cost report until final is true, or label the number as in progress.
Held money is also not spend. If you are reconciling against a budget, watch both debited and held, because held money can turn into captured spend or be refunded.
What should I do with the answer?
Use it to decide where to optimise. If one operation type dominates, change that step: pick a cheaper tier, shorten the clips, or move a repeated step into a Format. If several steps are close, the cost is spread out and the saving is in doing less overall.
For a run, the cap field shows the generation cap and what is left. It is a limit, not a cost, so do not report it as spend.
Finally, keep the call cheap. A usage read is a read, so it counts against your read budget rather than spending money, but polling it in a tight loop still burns request capacity. Read it once when the thread settles.
Sources
Related posts
More in Developers
- Sume API uptime monitor: use GET /v1/health, not /health
Point an uptime probe at https://api.sume.com/v1/health: no API key, returns apiVersion, status and build metadata. The unversioned /health is hidden.
- Veo 3.1 returns one video per request: how to get 4 variants
Google's Veo 3.1 table says one video per request. To get four variants on Sume, send four requests with four Idempotency-Keys, and watch queue_full.
- AI SDK stream cancel on disconnect: the Sume job keeps running
AI SDK 7.0.127 fixes stream cancellation when consumers disconnect. A cancelled stream does not cancel a Sume job: save the job id and read status later.
- AI SDK 7 tool search with deferred tools and Sume tools
AI SDK 7.0.127 lets a search() callback rank eligible deferred tools. Load a few Sume tools up front and fetch the rest by tools_schema on demand.
Written by Sume