Haiku 5.5: cached tokens count toward the 100k price step

Anthropic counts cache reads and writes in a Haiku 5.5 prompt's length, so a 90k cached plus 15k fresh turn is billed at the over-100k rates. Worked example.

4 min readSume
All posts

Short answer

Yes. Anthropic's pricing page says a Haiku 5.5 request's prompt length counts all of its input tokens, including cache reads and cache writes, and that a request over 100,000 tokens pays the higher prices even when part of the prompt is a cache hit. Each request is priced on its own.

For a long-running agent, that means a growing conversation can cross the line mid-session. The turn that crosses it costs more per token on everything, not only on the tokens past 100,000.

Worked example

Take a turn with 90,000 tokens read from cache, 15,000 fresh input tokens and 2,000 output tokens. The prompt is 105,000 tokens, so it is over the threshold.

One Haiku 5.5 request at both tiers (Anthropic pricing page, read 2026-10-08)
LineTokensUp to 100k rate $/MOver 100k rate $/MCost at over-100k
Cache read90,0000.010.0590,000 x 0.05 / 1M = $0.0045
Fresh input15,0000.100.5015,000 x 0.50 / 1M = $0.0075
Output2,0000.502.502,000 x 2.50 / 1M = $0.0050
Total107,000$0.0170

What the same turn would cost under 100k

If the prompt were 95,000 tokens (say 80,000 cached and 15,000 fresh), the same output would price at the base rates: 80,000 x 0.01 / 1M = $0.0008, 15,000 x 0.10 / 1M = $0.0015, 2,000 x 0.50 / 1M = $0.0010, total $0.0033. The over-100k version of the 105,000-token turn is about five times that, for a prompt that is about 10 percent longer.

Anthropic's batch table has the same step: $0.05 and $0.25 up to 100k, $0.25 and $1.25 above.

In a Sume run

Sume's registry rate card for Haiku 5.5 is the base tier (input $0.10, output $0.50, cache read $0.01, cache write $0.125). The repo comment says a Haiku 5.5 turn reports usage per request, so each request settles on its own tier, while reservations and estimates use the base card. Expect a long session to be billed higher than a base-tier estimate suggests. Trimming the prompt before it grows past 100,000 tokens is the lever.

Keeping a long session under the line

The step punishes growth you did not choose. A long conversation carries every earlier message, so the prompt creeps toward 100,000 tokens even if each new message is short. Tools listed in the prompt count too: Anthropic says the tools parameter and tool definitions add input tokens, and Haiku 5.5 adds 286 tokens of tool-use system prompt with automatic tool choice and 406 with any or a named tool.

The levers are the same ones that keep any agent cheap: trim old tool results, summarize the conversation before it crosses the line, and load only the tools a step needs. The cache does not help you cross the line cheaply, because Anthropic counts cached tokens in the length; it only discounts them within the tier.

  • Tool-use system prompt on Haiku 5.5: 286 tokens (auto, none), 406 (any, tool).
  • A request over 100,000 tokens is priced at 5x on every token class: input, output, cache read and cache write.

Sources

Related posts

More in Models

All Models posts

Written by Sume