Haiku 5.5 tool-use overhead: 286 tokens, 406 with tool_choice any

Claude Haiku 5.5 adds 286 system-prompt tokens when tools are present and 406 when tool_choice is any or tool. What that costs a Sume agent loop.

5 min readSume
All posts

When you send any tools to Claude Haiku 5.5, Anthropic adds a hidden tool-use system prompt of 286 tokens if tool_choice is auto or none, and 406 tokens if it is any or tool. That is on top of the tool definitions, the tool calls and the tool results. For a Sume agent loop at the lower prompt price of $0.10 per million input tokens, the overhead is a fraction of a cent per request, but it repeats on every turn of a polling loop.

The numbers

The page notes that if no tools are provided, a tool_choice of none adds 0 tokens. It also says the counts are in addition to the tokens from the tools parameter, tool_use blocks and tool_result blocks.

Tool-use system prompt tokens, Anthropic pricing page, read 2026-10-08
Modeltool_choice auto or nonetool_choice any or tool
Claude Haiku 5.5286406
Claude Haiku 4.5496588
Claude Sonnet 5.5286Not listed
Claude Opus 5.5286Not listed

Reading it for a Sume loop

The fixed part is small. The variable part is the tool definitions, which grow with the Sume tools you expose. With tool search, the loaded definitions count as input tokens like any other, and tool search itself is not metered separately.

Arithmetic from the published rate: 406 tokens at $0.10 per million input tokens is about $0.00004 per request. Even 100 polling turns add well under a cent. The number to watch is the size of tool results returned by jobs_wait and jobs_result, not the overhead.

Where the Sume side differs

Sume's hosted MCP server has a progressive tool surface in the repo, in which search_tools, get_tool_details and call_tool appear only on Studio Agent sandbox credentials, not on a human OAuth session or an API key. That is a presentation choice, not authorization: scopes and gates are the same. For your own client, the tools you see from tools_list are the ones whose definitions you pay to send.

Use tool_choice of auto unless you have a reason to force a call, since any and tool cost 120 more tokens per request. If you force a tool for a paid call, still send idempotency_key and max_spend_usd. See stable tool order and the prompt cache.

Count your own overhead

Anthropic's page says you can estimate token counts with the token counting endpoint. Send the same request with the full Sume tool list and with a trimmed one, and compare the totals. That measures your real tool-definition cost, which the 286 and 406 figures do not include.

Then decide with arithmetic. If the trimmed list saves a few thousand tokens per request over hundreds of turns, it is worth the work. If the difference is a few hundred tokens, spend the effort on result sizes instead.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume