Claude Code token usage vs Sume spend: two meters, one session

Claude Code's token totals and Sume's generation spend are billed in different places. Log tokens with a mod, read Sume's wallet with usage_get, and compare.

5 min readSume
All posts

They are two separate meters. Claude Code's token usage is what the model consumed on your Claude plan or API key, and a mod can read it from the turn.complete and turn.step events. Sume's generation spend is what hosted jobs such as generate_image or generate_video drew from your Sume wallet, and it shows up in usage_get and balance_get, or sume usage get and sume balance in the CLI. A turn that wrote a prompt and submitted a clip appears in both, in different units.

The event fields come from Claude Code's events guide, read 2026-10-03; the Sume tools come from tools and gates and the CLI reference.

What can a mod see about tokens?

Claude Code's guide says turn.complete fires when a turn ends, including one you interrupted, and that e.usage holds the turn's token totals while e.durationMs is how long it took. turn.step fires before each request to the model, and its result's usage has input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens and the model that answered. A subagent's turn fires the same events with e.agentId set, so you can separate it from the main conversation. This mod logs one line per turn without changing anything:

export function register(on) {
  on('turn.complete', async ($, e, next) => {
    $.ui.log('turn tokens ' + JSON.stringify(e.usage) + ' in ' + e.durationMs + ' ms')
    return next(e)
  })
}

What can it not see?

Sume spend. The mod sees tool calls, and a Sume tool returns a job, not a price in the turn's token totals. To read the wallet side, call Sume's usage_get and balance_get tools from the agent, or run the CLI after sume login:

The reverse also holds: Sume does not know how many tokens the agent used. Its records are about jobs, so a session that burned many tokens and made no paid call leaves nothing in usage_get. When you reconcile a session, expect one meter to move without the other.

sume balance
sume usage get --limit 20

How do the two meters compare?

The table puts the two side by side. The rows that matter most are who bills it and whether you can preview it before the spend happens.

Where each kind of spend is measured, from Claude Code and Sume docs, read 2026-10-03.
Claude Code tokensSume generation spend
What it measuresTokens the model used for the turnCredits and wallet drawn by hosted jobs
Where to read itturn.complete e.usage in a mod, turn.step usage per requestusage_get, balance_get, sume usage get, sume balance
Who bills itYour Claude plan or API keyYour Sume wallet
Preview before spendingNot in a moddry_run=true or generation_admission_preview
CapYour plan's or key's own limitsmax_spend_usd per call, only when sent; wallet balance
Subagent splite.agentId on eventsNot separated by Claude Code subagent

Why does a long agent loop surprise people?

Because the loop raises the token meter on every step while paid tools raise the Sume meter on some of them. A prompt that retries a failed clip five times costs five generations plus five rounds of conversation. Sume's idempotency_key keeps a retried identical call from creating a duplicate job, and jobs_wait should be called again with the same ids rather than resubmitting. Neither changes the token meter, so look at both numbers after a long session.

Subagents widen the gap. Claude Code fires turn events for a subagent with its own e.agentId, so the token side can be split by agent, while Sume sees only the account and key that made each call. If you need per-agent spend on Sume's side, give each unattended job its own bounded call and record the job it returns.

How do I wire a check?

You do not need a dashboard to catch a runaway session. A mod, a CLI command and one Sume argument cover most of it.

  • Log tokens per turn with the mod above, so a runaway turn is visible in the transcript.
  • At the end of a session, run sume usage get --limit 20 or ask the agent for usage_get, and compare the entries to the turns you expected.
  • Set max_spend_usd on paid calls and use dry_run=true first for bursts; Sume enforces the cap only when it is sent.
  • For unattended runs, use Agent Completions, which require a generation_spend_cap_usd on each call, rather than a bare MCP session.
  • Remember that a mod may not be loaded (--safe-mode, a crash loop, or disableAllHooks), so the log is a convenience, not a ledger.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume