Haiku 5.5 prompt caching for a Sume tool list: what breaks the cache
Haiku 5.5 cache hits cost $0.01 per million tokens. Keep the Sume tool list and effort setting stable so a long agent run keeps hitting the cache.

To get cache hits on Claude Haiku 5.5 in a loop that calls Sume tools, keep three things stable between requests: the order and content of the tool definitions, the system prompt, and the top-level effort value. On the pricing page read on 2026-10-08 a Haiku 5.5 cache hit costs $0.01 per million tokens for prompts up to 100,000 tokens, against $0.10 for uncached input. Anthropic's effort page says that changing top-level effort between requests does not preserve cached prefixes from earlier turns.
What the vendor pages say
The pricing page says a 5-minute cache pays off after one read, and the 1-hour cache after two.
| Item | Fact |
|---|---|
| Cache hit (prompt up to 100k) | $0.01 per million tokens |
| 5-minute cache write | $0.125 per million tokens (1.25x base) |
| 1-hour cache write | $0.20 per million tokens (2x base) |
| Cache hit multiplier | 0.1x base input |
| Top-level effort changed | Does not keep earlier cached prefixes |
| Deferred tool with cache_control | 400 error; put the breakpoint on a non-deferred tool |
The tool list is the long prefix
In an agent that uses Sume's hosted MCP server, the biggest stable prefix is the tool list. Sume's registry returns tools in a fixed order, which helps; if your client re-sorts or filters the list differently on each turn, the prefix changes and the cache misses. The same holds if you add and remove tools mid-run. Decide the visible set at the start from the scope you have: an API key or an OAuth session with mcp:write shows the full tool set, and a read-only session shows fewer.
If you use Anthropic's tool search with deferred loading, the deferred definitions are excluded from the prefix, and Anthropic says the prefix is untouched when a tool is discovered, so caching is preserved. Put cache_control on a non-deferred tool.
Practical rules
The Sume side has its own note on stable tool order and the prompt cache.
- Fix effort at the start of a cached run, or use the per-message effort beta instead of changing the top-level value.
- Keep volatile text such as timestamps and job ids out of the system prompt; put them in the user turn.
- Use a 1-hour write for a run that polls for longer than five minutes between turns, if the pricing math works for your volume.
- Check the cache-read count in the usage block rather than assuming a hit.
How to confirm the cache is working
Look at the usage block on each response. Anthropic's pricing page defines cache writes and cache reads as separate charges, so a healthy loop shows a write on the first turn and reads on the turns after it. If every turn shows a write, something in the prefix changes between requests.
The usual suspects in a Sume loop are a tool list that is rebuilt in a different order, a timestamp in the system prompt, a changed effort value, and a tool added after a scope upgrade. Fix the first one you find, then check again.
Sources
Related posts
More in Developers
- Haiku 5.5 returns 400 for thinking disabled at xhigh: a Sume agent fix
Claude Haiku 5.5 rejects thinking disabled at xhigh or max effort. How that 400 shows up in an agent that calls Sume, and how to set the pair.
- Hono 4.13.10 split adapters: update a Sume webhook receiver
Hono 4.13.10 moved adapters to @hono/bun, @hono/deno and @hono/cloudflare-workers. The Sume verifyWebhook call needs no change; only the entrypoint imports do.
- Hono 4.13.13 deprecates app.mount(): mount a Sume webhook sub-app
Hono 4.13.13 deprecates app.mount() in favor of Mount Middleware. A Sume webhook receiver written as a Hono sub-app uses app.route() and keeps its raw body.
- Why input_references fails on Imagen, Recraft, Soul and Qwen Max
Five Sume image rows list zero input_references: imagen-4 fast and ultra, recraft-v4, Soul and qwen-image-max. Check the catalog in Python first.
Written by Sume