One agent turn's total cost: pass its job id to /v1/usage

Pass an agent turn's job id as job_id on GET /v1/usage and Sume sums the turn's own model row plus every job it commissioned, with turn_job_id on each row.

5 min readSume
All posts

To get the total cost of one agent turn, call GET /v1/usage?job_id=<turn job id>. The job_id filter accepts a generation job id and also the job id of a turn; with a turn id, the summary includes the turn's own model row and every job that the turn commissioned.

This is stated on the Usage page (read 2026-10-11) and checked by the API's usage-scope test. It sits between two other scopes: thread_id for a whole conversation and run_id for one Format, Action or Agent run. A turn is the middle step, and it is the one that answers "what did this single request to the agent cost".

Which rows does a turn scope pick up?

The scope selects rows two ways. A row matches when its own job_id equals the id you passed. It also matches when the turn that commissioned it, stamped on the reservation as the agent job id, equals the id you passed. The first arm finds the turn's own studio_agent_llm row; the second finds the generation jobs, speech jobs and script children it started.

Each public row then carries turn_job_id: the agent turn that commissioned it, or null for rows outside a turn. In the API test, the turn's own LLM row has turn_job_id: null, because it is the turn, while its children carry the turn's id. So filtering your own copies of the rows on turn_job_id finds the children, not the turn row.

Three usage scopes, from the Usage page and API usage-scope code (read 2026-10-11)
Query paramScopeIncludes
thread_idOne agent threadEvery turn, sidecar, Browser session and job in the thread; also split by run
run_idOne Format, Action or Agent runThe run's rows, plus its generation cap accounting
job_id (generation job)One jobThat job's ledger rows
job_id (turn)One agent turnThe turn's model row and every job it commissioned

Why not just use the run receipt?

A run receipt's billable_amount_usd_micros is the generation cap's figure: reserved plus captured generation rows, without the model's own turns. The Usage page says outright that this value is never a cost. The scoped summary's debited_usd_micros is the figure to quote, because it counts captured rows of every operation type, the agent's own turns included.

A turn scope gives you that same honest figure at a finer grain. If a thread had six turns and you want to know which one made the expensive video, query each turn's job_id, or read by_operation_type on the thread, as in which step cost most in a thread.

What can go wrong

Turns end before their jobs do. A video the turn commissioned may still be processing, so the summary shows final: false and a hold in held_usd_micros. Quote the total only once final is true, as in final false means a hold is still open.

Rows are stamped at reservation time. Rows written before the stamps existed are found by other columns on the job, so an old thread can read differently from a new one. And as with every scope, limit caps the listed rows only; the summary still folds every row the scope caused.

A request that prints the turn total

The response has the same shape as a thread read: data.usage lists rows, and data.summary carries the totals.

import os
import requests


def main():
    r = requests.get(
        "https://api.sume.com/v1/usage",
        params={"job_id": os.environ["TURN_JOB_ID"], "limit": 50},
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
        timeout=30,
    )
    r.raise_for_status()
    data = r.json()["data"]
    s = data["summary"]
    print("turn cost USD:", s["debited_usd"], "final:", s["final"])
    for row in data["usage"]:
        print(row["operation_type"], row["status"],
              row["billable_amount_usd_micros"], row["turn_job_id"])


main()

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume