Sonnet 5.5 token bill vs Sume spend cap: two meters on one agent run

Sonnet 5.5 tokens bill at the model vendor. Sume generation bills in the Sume wallet. A worked example shows why one cap cannot cover both.

5 min readSume
All posts

An agent run that uses Sonnet 5.5 and calls Sume has two separate bills. Anthropic lists Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 and cache writes at $2.50 (read 2026-10-05). That meter covers model tokens. Sume's generation_spend_cap_usd caps media generation in the Sume wallet. Neither cap covers the other meter.

A worked example, with invented token counts

The token counts below are made up to show the arithmetic. They are not a measurement of any real run. Suppose an agent loop sends 200,000 input tokens and 20,000 output tokens to Sonnet 5.5.

Input: 200,000 / 1,000,000 x $2 = $0.40. Output: 20,000 / 1,000,000 x $10 = $0.20. The token total is $0.60. Now suppose the same loop starts one Sume run with a $3 generation cap. The most you could be charged across both meters is $3.60, but only the $3 belongs to Sume.

Two meters for one illustrative run, read 2026-10-05
MeterRate or limitIllustrative amount
Sonnet 5.5 input$2 per 1M tokens200,000 tokens = $0.40
Sonnet 5.5 output$10 per 1M tokens20,000 tokens = $0.20
Sonnet 5.5 cache read$0.20 per 1M tokensBilled apart from the above
Sume generation capgeneration_spend_cap_usd in the request$3.00 ceiling
Total ceiling in this exampleToken bill plus Sume cap$3.60

What each cap can and cannot do

Sume documents that usage.billable_amount_usd_micros on a Format or Action receipt is the generation spend that the run's cap is enforced against, and that it does not include the agent's own LLM turn. Do not read the cap as a total cost for the run.

The reverse also holds. A token budget on the Anthropic side does not stop a Sume job that was already accepted. Cancel works only before generation starts, after which the job completes or fails normally.

Compute both in code

A small function keeps the two meters apart in your own logs. It runs offline.

def token_usd(inp: int, out: int, in_rate: float = 2.0,
              out_rate: float = 10.0) -> float:
    return inp / 1_000_000 * in_rate + out / 1_000_000 * out_rate

def ceiling_usd(inp: int, out: int, sume_cap: float) -> float:
    return round(token_usd(inp, out) + sume_cap, 2)

if __name__ == "__main__":
    print(round(token_usd(200_000, 20_000), 2))
    print(ceiling_usd(200_000, 20_000, 3.0))

Why the cap is not a token budget

It is tempting to size the Sume cap from the token bill, because both are dollar amounts. They answer different questions. Tokens scale with how much the model reads and writes. Generation spend scales with how many clips, images or voice minutes the run asks for, and with the model Sume routes each request to.

So size the Sume cap from the rate card for the media in the task. A single 30 second clip and a thirty-image batch have wildly different generation cost for the same number of tokens in the prompt. Use the pricing page for the media types in play and add headroom for one retry, then let the cap enforce it.

Practical rules

Keep the two numbers in separate fields in your dashboards.

  • Set the Sume cap per task, from the Sume pricing page, not from the token budget.
  • Log token cost and Sume spend separately, and sum them only for reporting.
  • Use GET /v1/usage as the authoritative Sume record, as the run webhook docs advise.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume