Haiku 5.5 charges more above 100k tokens: keep Sume tools lean

Claude Haiku 5.5 prices prompts over 100,000 tokens at five times the base rate. How a long Sume tool list and job history can push you over.

6 min readSume
All posts

Claude Haiku 5.5 is priced by prompt length, and a prompt over 100,000 tokens is billed at five times the rate of a shorter one. At the shorter rate input is $0.10 per million tokens and output is $0.50; above 100,000 tokens they are $0.50 and $2.50. For an agent that talks to Sume's hosted MCP server, that makes the length of the tool list and of the accumulated job results a cost question, not only a context question.

The two price rows

Anthropic's pricing page lists Haiku 5.5 as two rows. The page also says the higher prices apply to a prompt over 100,000 tokens, and that other 4.6-and-later models include the full 1M context window at standard pricing, with Haiku 5.5 the stated exception.

Claude Haiku 5.5 price per million tokens, Anthropic pricing page, read 2026-10-08
ItemPrompt up to 100,000 tokensPrompt over 100,000 tokens
Base input$0.10$0.50
5-minute cache write$0.125$0.625
1-hour cache write$0.20$1
Cache hit$0.01$0.05
Output$0.50$2.50
Batch input / output$0.05 / $0.25$0.25 / $1.25

What grows the prompt in a Sume loop

Three things accumulate. The first is the tool definitions, which are sent on every request; Anthropic's tool-search page puts a typical multi-server setup at about 55k tokens of definitions before any work. The second is tool results: jobs_result and jobs_wait with results included can return large payloads. The third is the transcript of earlier polls.

Sume documents tools_list and tools_schema as the discovery pair: list what the session can see, then fetch one contract by name. Fetching one schema at the moment you need it keeps the other tool contracts out of the prompt.

Ways to stay under the line

None of this changes what Sume charges. Sume's generation spend is a separate meter from model tokens, and max_spend_usd or generation_spend_cap_usd caps only the Sume side.

  • Expose only the Sume tools the task needs. A read-only OAuth session sees only the read-only tools, which is also a smaller prompt.
  • Return job ids and statuses to the model, and read full results only for the job that finished.
  • Keep tool order stable so a cached prefix is reused; see stable tool order and the prompt cache.
  • Start a fresh thread for a new render instead of carrying a long polling history.

A quick check

Read the input token count your client reports at the end of a typical run. If it sits near 80,000, the next added tool or result can move the whole prompt to the $0.50 row, and the fix is cheaper to apply before that happens. The Claude Code side of the same problem is covered in keeping the tool list small on a 1M context.

A worked comparison from the published rates

Take one million input tokens and one million output tokens in a single long session. At the shorter-prompt rows that is $0.10 plus $0.50, or $0.60. At the over-100,000-token rows it is $0.50 plus $2.50, or $3.00. The figures come straight from Anthropic's table; they are not a measurement of any Sume run.

The practical point is the step, not the total. A session that stays under the line keeps the lower rows, and one that crosses it pays the higher rows for the prompts that exceed it. Anthropic's wording is that a prompt of over 100,000 tokens pays higher prices, so check your own bill for how the boundary is applied before you build a budget on it.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume