claude -p total_cost_usd vs your Sume bill: which number to trust
claude -p prints total_cost_usd for the Claude side of a run. Sume's generation spend is billed on Sume and read from GET /v1/usage. Here is how to log both.

total_cost_usd in the claude -p --output-format json result is Claude Code's own figure for the model usage in that session. It is not your Sume bill. Generation that your agent starts through the Sume hosted MCP server is billed by Sume, and GET /v1/usage is the record for it, so a script that tracks spend needs to read both.
Two meters, two owners
Claude Code's headless docs say that with --output-format json the payload includes total_cost_usd and a per-model breakdown, so a scripted caller can track spend without opening a dashboard. That is useful for the model side of an unattended run. It cannot see what the Sume tools charged, because those charges are made on Sume's servers when a generation job is admitted.
What Sume counts
Sume's own receipts follow the same split. On Format runs and scheduled runs, usage.billable_amount_usd_micros covers generation spend only and leaves out the agent's LLM turn. The docs call GET /v1/usage the authoritative billing record, and a figure on a receipt is a receipt value, not an invoice.
| Question | Number to read | Where it comes from |
|---|---|---|
| What did the Claude side of this script cost? | total_cost_usd | claude -p JSON output, computed by Claude Code |
| What did generation cost for this run? | usage.billable_amount_usd_micros | Sume receipt, generation only |
| What was I actually billed? | GET /v1/usage | Sume usage ledger, authoritative |
| How much can this run spend at most? | generation_spend_cap_usd or max_spend_usd | Your request, enforced by Sume |
Log both numbers after a headless run
The shell script below runs one tiny headless call against the hosted server, then prints both numbers. It builds the MCP config with jq so your key never lands in a file you commit, and it assumes claude, jq and curl are installed and SUME_API_KEY is set. The prompt only calls mcp_health, a read tool, so it generates nothing and spends nothing on Sume.
jq -n --arg k "$SUME_API_KEY" '{mcpServers:{sume:{type:"http",url:"https://mcp.sume.com/mcp",headers:{Authorization:("Bearer "+$k)}}}}' > /tmp/sume-mcp.json
claude -p "Call the mcp_health tool and print the result." \
--mcp-config /tmp/sume-mcp.json --output-format json > /tmp/run.json
echo "claude side:"
jq -r '.total_cost_usd' /tmp/run.json
echo "sume side:"
curl -sS https://api.sume.com/v1/usage \
-H "Authorization: Bearer $SUME_API_KEY" | head -c 600Where to put the limit
Run the script once with a read-only prompt to see the first number, then once after a capped paid call to see the second move. Keep the two separate in your logs. If you add them, label the sum as an estimate.
For paid calls, do not rely on either meter to stop a loop. Put the limit on the request: max_spend_usd on an MCP paid tool, generation_spend_cap_usd on a Format run, and an idempotency_key on every create so a retry does not bill twice.
Short version
The two numbers answer different questions, and each is accurate for its own question. Read the first for the Claude side, read the second for what Sume charged, and trust the ledger when they disagree with a receipt.
Sources
Related posts
More in Developers
- Sonnet 5.5 token bill vs Sume spend cap: two meters on one agent run
Sonnet 5.5 tokens bill at the model vendor. Sume generation bills in the Sume wallet. A worked example shows why one cap cannot cover both.
- Cloudflare Worker for a social video pipeline: SDK verifyWebhook
@sume-com/sdk 0.2.0 needs only fetch and WebCrypto, so verifyWebhook runs in a Cloudflare Worker. Verify the raw body, then answer 2xx before you post.
- Cloudflare Worker that verifies a Sume video webhook (verifyWebhook)
A 20-line Cloudflare Worker for Sume job webhooks: verifyWebhook from @sume-com/sdk on WebCrypto, empty-secret guard, 204 fast ack. Why the check is async.
- compose_duration_clamped_to_source: the banner clip came out shorter
Compose clamps video.duration to the source file and warns compose_duration_clamped_to_source. The job still succeeds; read the result length first.
Written by Sume