Haiku 5.5: cached tokens count toward the 100k price step
Anthropic counts cache reads and writes in a Haiku 5.5 prompt's length, so a 90k cached plus 15k fresh turn is billed at the over-100k rates. Worked example.

Short answer
Yes. Anthropic's pricing page says a Haiku 5.5 request's prompt length counts all of its input tokens, including cache reads and cache writes, and that a request over 100,000 tokens pays the higher prices even when part of the prompt is a cache hit. Each request is priced on its own.
For a long-running agent, that means a growing conversation can cross the line mid-session. The turn that crosses it costs more per token on everything, not only on the tokens past 100,000.
Worked example
Take a turn with 90,000 tokens read from cache, 15,000 fresh input tokens and 2,000 output tokens. The prompt is 105,000 tokens, so it is over the threshold.
| Line | Tokens | Up to 100k rate $/M | Over 100k rate $/M | Cost at over-100k |
|---|---|---|---|---|
| Cache read | 90,000 | 0.01 | 0.05 | 90,000 x 0.05 / 1M = $0.0045 |
| Fresh input | 15,000 | 0.10 | 0.50 | 15,000 x 0.50 / 1M = $0.0075 |
| Output | 2,000 | 0.50 | 2.50 | 2,000 x 2.50 / 1M = $0.0050 |
| Total | 107,000 | $0.0170 |
What the same turn would cost under 100k
If the prompt were 95,000 tokens (say 80,000 cached and 15,000 fresh), the same output would price at the base rates: 80,000 x 0.01 / 1M = $0.0008, 15,000 x 0.10 / 1M = $0.0015, 2,000 x 0.50 / 1M = $0.0010, total $0.0033. The over-100k version of the 105,000-token turn is about five times that, for a prompt that is about 10 percent longer.
Anthropic's batch table has the same step: $0.05 and $0.25 up to 100k, $0.25 and $1.25 above.
In a Sume run
Sume's registry rate card for Haiku 5.5 is the base tier (input $0.10, output $0.50, cache read $0.01, cache write $0.125). The repo comment says a Haiku 5.5 turn reports usage per request, so each request settles on its own tier, while reservations and estimates use the base card. Expect a long session to be billed higher than a base-tier estimate suggests. Trimming the prompt before it grows past 100,000 tokens is the lever.
Keeping a long session under the line
The step punishes growth you did not choose. A long conversation carries every earlier message, so the prompt creeps toward 100,000 tokens even if each new message is short. Tools listed in the prompt count too: Anthropic says the tools parameter and tool definitions add input tokens, and Haiku 5.5 adds 286 tokens of tool-use system prompt with automatic tool choice and 406 with any or a named tool.
The levers are the same ones that keep any agent cheap: trim old tool results, summarize the conversation before it crosses the line, and load only the tools a step needs. The cache does not help you cross the line cheaply, because Anthropic counts cached tokens in the length; it only discounts them within the tier.
- Tool-use system prompt on Haiku 5.5: 286 tokens (auto, none), 406 (any, tool).
- A request over 100,000 tokens is priced at 5x on every token class: input, output, cache read and cache write.
Sources
Related posts
More in Models
- Can a Haiku 5.5 agent look at images? Sume's 30-image attachment rules
Anthropic says all current Claude models take image input. Sume accepts up to 30 images, 30 MB each, 500 MB total; Haiku 5.5 is a Format-run pick.
- 1,000 agent turns: Haiku 5.5 about $4, GPT-6 Astra about $400
At list rates, 1,000 turns of 30,000 input and 2,000 output tokens cost about $4 on Haiku 5.5 and $400 on GPT-6 Astra. The arithmetic and the cache effect.
- Higgsfield Soul on Sume: half a cent per image, batches of 1 or 4
higgsfield-soul costs $0.005 at 720p and $0.0075 at 1080p on Sume. It takes batch sizes 1 or 4, no references. Price table and a four-image call.
- Higgsfield Soul on Sume: 1 cent an image, n of 1 or 4, $100 per 10,000
Higgsfield Soul is the cheapest Sume image model at 1 cent. It is text-only, takes n of 1 or 4 and 720p or 1080p, so 10,000 images is 2,500 calls and $100.
Written by Sume