Usage hold_counts: open, browser_session, billing_pending
A scoped Sume usage summary can show final false with money held. hold_counts says whether a turn runs, a Browser session is open, or work is being priced.

When a scoped usage summary shows final: false, read hold_counts next: it splits the reserved rows into open (a turn or generation job is still running), browser_session (a Browser session is still open) and billing_pending (the work ended and Sume is still pricing it). Only open means something is still being produced; the other two are money waiting to settle.
The summary comes from GET /v1/usage when you add thread_id, run_id or job_id, as described on the Usage page (read 2026-10-11). The three hold_counts fields are defined in the API's usage schema and exercised in its tests, and they always sum to rows.reserved. If you already know that final means no hold is open, this post is the next step: telling the holds apart.
What are the three buckets?
A hold is a reserved ledger row: Sume reserved the estimated usage before the work ran, as the status table on the Usage page says. Captured rows are spend, refunded rows are given back, and reserved rows are neither yet. The summary reports their total as held_usd_micros and states plainly that it is not spend yet.
hold_counts then classifies each reserved row by what it is waiting on. The classification uses the row's operation type and its settle state, so it does not guess from age.
| Bucket | What the reserved row is | When it settles | Still producing work? |
|---|---|---|---|
| open | An agent turn's model ceiling, or a generation job in flight | When the turn or job completes, fails or is canceled | Yes |
| browser_session | A Browser session's ceiling | When the session ends (idle end or its timeout) | No, but the session is open |
| billing_pending | Work that ended; the final usage is not recorded yet | When the settle sweeper prices the row | No |
Why the Browser bucket is separate
A Browser session can outlive the turn that opened it. If a thread's last turn finished but the session is still inside its idle window, the thread has a reserved Browser row and nothing else running. Counting that as an open turn hold would tell you the agent is still working when it is not.
So the summary keeps it apart. The Browser row settles when the session ends, which means final stays false until then and debited_usd_micros can still rise by the Browser charge. A cost dashboard that prints a total should therefore wait for final: true, or label the figure as provisional while browser_session is above zero.
What billing_pending means for a stopped turn
The other settle-state case is a turn that was stopped. Its last answer used model tokens, and the usage for that answer reaches the ledger a moment later. Until it does, the row is parked with a pending_ settle state (or unpriced), and the summary counts it as billing_pending, not as an open hold.
The practical reading: do nothing. Do not retry the turn, and do not read the figure as a leak. The summary also reports pending_usd_micros, the total of rows parked this way, so you can see how much of held_usd_micros is only waiting to be priced.
A check that tells the three apart
This script prints the figures and decides whether the total can be quoted. It reads the summary only; it never adds rows itself, which the Usage page warns against.
import os
import requests
def main():
r = requests.get(
"https://api.sume.com/v1/usage",
params={"thread_id": os.environ["THREAD_ID"], "limit": 50},
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
s = r.json()["data"]["summary"]
h = s["hold_counts"]
print("debited_usd:", s["debited_usd"], "final:", s["final"])
if h["open"]:
print("work still running:", h["open"], "hold(s)")
elif h["browser_session"]:
print("Browser session open; total can still rise")
elif h["billing_pending"]:
print("priced shortly; recheck")
else:
print("safe to quote")
main()
Where this fits in a spend report
Quote debited_usd from the summary and gate it on final. Keep held_usd_micros out of any cost figure, because a hold is a ceiling, not a charge, and a refunded hold goes back to the balance. For a per-step split of the same thread, see which step cost most in a thread, and for the closing rule see final false means a hold is still open.
Nothing here changes what you pay. It changes how long you should wait before treating a number as the bill.
Sources
Related posts
More in Developers
- What one script_run call cost: the usage script_runs array
Sume's usage summary lists each script_run call with its rows, debited, held and refunded micros, so a fan-out of TTS or image jobs has one price.
- WireMock scenarios: fake a Sume job going queued to completed
Use one WireMock scenario per job so GET /v1/jobs/{id}/status answers queued, processing, then completed, and test the poll loop without spending.
- Word timestamps for an Eleven v4 voice: the router says 400, use STT
The Sume TTS Router rejects timestamps on eleven-* ids. Get word timings by running the finished narration through Sume STT for one cent a minute.
- xAI video API polling: pending, done, failed vs Sume job states
xAI's video API returns a request_id and three states. Here is how they map onto Sume's five /v1/videos statuses, with a runnable Python polling loop.
Written by Sume