Claude Code token usage vs Sume spend: two meters, one session
Claude Code's token totals and Sume's generation spend are billed in different places. Log tokens with a mod, read Sume's wallet with usage_get, and compare.

They are two separate meters. Claude Code's token usage is what the model consumed on your Claude plan or API key, and a mod can read it from the turn.complete and turn.step events. Sume's generation spend is what hosted jobs such as generate_image or generate_video drew from your Sume wallet, and it shows up in usage_get and balance_get, or sume usage get and sume balance in the CLI. A turn that wrote a prompt and submitted a clip appears in both, in different units.
The event fields come from Claude Code's events guide, read 2026-10-03; the Sume tools come from tools and gates and the CLI reference.
What can a mod see about tokens?
Claude Code's guide says turn.complete fires when a turn ends, including one you interrupted, and that e.usage holds the turn's token totals while e.durationMs is how long it took. turn.step fires before each request to the model, and its result's usage has input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens and the model that answered. A subagent's turn fires the same events with e.agentId set, so you can separate it from the main conversation. This mod logs one line per turn without changing anything:
export function register(on) {
on('turn.complete', async ($, e, next) => {
$.ui.log('turn tokens ' + JSON.stringify(e.usage) + ' in ' + e.durationMs + ' ms')
return next(e)
})
}What can it not see?
Sume spend. The mod sees tool calls, and a Sume tool returns a job, not a price in the turn's token totals. To read the wallet side, call Sume's usage_get and balance_get tools from the agent, or run the CLI after sume login:
The reverse also holds: Sume does not know how many tokens the agent used. Its records are about jobs, so a session that burned many tokens and made no paid call leaves nothing in usage_get. When you reconcile a session, expect one meter to move without the other.
sume balance
sume usage get --limit 20How do the two meters compare?
The table puts the two side by side. The rows that matter most are who bills it and whether you can preview it before the spend happens.
| Claude Code tokens | Sume generation spend | |
|---|---|---|
| What it measures | Tokens the model used for the turn | Credits and wallet drawn by hosted jobs |
| Where to read it | turn.complete e.usage in a mod, turn.step usage per request | usage_get, balance_get, sume usage get, sume balance |
| Who bills it | Your Claude plan or API key | Your Sume wallet |
| Preview before spending | Not in a mod | dry_run=true or generation_admission_preview |
| Cap | Your plan's or key's own limits | max_spend_usd per call, only when sent; wallet balance |
| Subagent split | e.agentId on events | Not separated by Claude Code subagent |
Why does a long agent loop surprise people?
Because the loop raises the token meter on every step while paid tools raise the Sume meter on some of them. A prompt that retries a failed clip five times costs five generations plus five rounds of conversation. Sume's idempotency_key keeps a retried identical call from creating a duplicate job, and jobs_wait should be called again with the same ids rather than resubmitting. Neither changes the token meter, so look at both numbers after a long session.
Subagents widen the gap. Claude Code fires turn events for a subagent with its own e.agentId, so the token side can be split by agent, while Sume sees only the account and key that made each call. If you need per-agent spend on Sume's side, give each unattended job its own bounded call and record the job it returns.
How do I wire a check?
You do not need a dashboard to catch a runaway session. A mod, a CLI command and one Sume argument cover most of it.
- Log tokens per turn with the mod above, so a runaway turn is visible in the transcript.
- At the end of a session, run
sume usage get --limit 20or ask the agent forusage_get, and compare the entries to the turns you expected. - Set
max_spend_usdon paid calls and usedry_run=truefirst for bursts; Sume enforces the cap only when it is sent. - For unattended runs, use Agent Completions, which require a
generation_spend_cap_usdon each call, rather than a bare MCP session. - Remember that a mod may not be loaded (
--safe-mode, a crash loop, ordisableAllHooks), so the log is a convenience, not a ledger.
Sources
Related posts
More in Developers
- claude plugin validate: read the hooks and calls lines of a mod
Before installing a Claude Code mod, run claude plugin validate and read its hooks and calls lines. Which combinations matter when Sume's MCP is connected.
- Contract-test Sume API responses against openapi.json (pytest)
Validate recorded Sume responses against the OpenAPI schema with jsonschema, including the OpenAPI 3.0 nullable fix. A tested pytest file and fixtures guide.
- DBOS Python durable workflow for a Sume job: resume after a crash
Submit and poll a Sume image job in a DBOS workflow: step retries, order-derived Idempotency-Key and workflow id, tested with DBOS 3.2.0 on SQLite.
- Dub one Short into 8 languages: Python fan-out and the total cost
Detach and transcribe once, then run one TTS job and one render per language. A Python fan-out and the per-Short bill, from Sume's catalog rates.
Written by Sume