Haiku 5.5 cache write vs read: break-even on a Sume tool catalog
Haiku 5.5 charges $0.125 per million to write cache and $0.01 to read it. For a 12,000-token Sume tool list the cache pays back on the first reuse.

Caching a Sume tool catalog in Claude Haiku 5.5 pays for itself on the first reuse. Anthropic lists cache writes at $0.125 per million tokens, cache reads at $0.01, and plain input at $0.10 (prompts up to 100,000 tokens). Writing costs $0.025 per million more than normal input, and each read saves $0.09 per million, so one reread more than repays the write premium.
The numbers below use one assumption you should replace with your own count: a tool block of 12,000 tokens that your client resends on every turn. Sume does not publish a token count for its tool list. You can see the size of yours by calling tools_list once through the hosted MCP endpoint and counting what your client puts in the prompt. Whether your client actually uses prompt caching is a client setting, not something the Sume server controls.
The price points, read 2026-10-09
These are the Haiku 5.5 rates for prompts of 100,000 tokens or fewer. Above that size Anthropic lists a second tier of $0.50 input, $0.625 cache write and $0.05 cache read per million.
| Item | Per million tokens (up to 100k prompt) | Per million tokens (over 100k prompt) |
|---|---|---|
| Input | $0.10 | $0.50 |
| Cache write | $0.125 | $0.625 |
| Cache read | $0.01 | $0.05 |
| Output | $0.50 | $2.50 |
The arithmetic for a 30-turn session
Assume 30 turns, each carrying the same 12,000-token tool block, and nothing else priced here. Without caching, every turn pays input rate. With caching, turn 1 writes the block and turns 2 to 30 read it.
- No cache: 30 x 12,000 = 360,000 tokens x $0.10 per million = $0.036.
- With cache: 12,000 x $0.125 per million = $0.0015 for the write, plus 29 x 12,000 = 348,000 tokens x $0.01 per million = $0.00348 for the reads. Total $0.00498.
- Saving: $0.036 - $0.00498 = $0.03102, about 86% of the tool-block cost.
- Break-even: the write premium is 12,000 x ($0.125 - $0.10) per million = $0.0003. One read saves 12,000 x $0.09 per million = $0.00108. The first reread already covers the premium.
What this does and does not change on the Sume side
The saving is on the model bill, not the generation bill. A hosted MCP session is billed for generation separately: Sume bills provider list price times 1.25, and a Wan 3.0 clip at 720p is $0.125 per second on the Sume catalog, so a five-second clip is $0.625. The $0.031 you save on 30 turns is small next to one clip, which is why the better lever on spend is the guardrails in the tools and gates page: dry_run, the optional max_spend_usd, and idempotency_key on every paid call.
Two practical notes. First, the tool block only stays cacheable if your client keeps it byte-identical between turns, so avoid reordering tools or injecting timestamps ahead of it. Second, if a turn needs three or more independent paid calls of the same shape, script_run collapses them into one tool result, which shrinks the turn count and therefore the repeated-prompt bill.
Sources
Related posts
More in Models
- HappyHorse 1.1 image-to-video: 300 px, 20 MB and ratio 1:2.5 to 2.5:1
Alibaba's HappyHorse 1.1 image-to-video accepts JPEG, PNG or WEBP of at least 300 px, up to 20 MB. Prepare a first frame, or use an Omni image_url on Sume.
- HappyHorse 1.1 API watermark defaults to on: how to turn it off
Alibaba's HappyHorse 1.1 image-to-video API adds a Happy Horse text mark bottom-right unless you set watermark to false. What to check before delivery.
- How many jobs make a 3-minute vertical video: 6, 12 or 18 by model
A 180 s vertical video needs 6 jobs on Wan 3.0 or Seedance 2.5, 12 on MiniMax H3 and 18 on Omni Flash 1.1, since each model has a different maximum clip length.
- 10-second Gemini Omni Flash 1.1 clip on Sume: $0.38 to $3.75
A 10-second Gemini Omni Flash 1.1 clip on Sume costs $0.38 at 360p up to $3.75 at 4K. The price is reserved at submit and refunded if the job fails.
Written by Sume