ENABLE_TOOL_SEARCH=false in Claude Code loads every Sume tool up front
Setting ENABLE_TOOL_SEARCH=false or a custom ANTHROPIC_BASE_URL turns off Claude Code tool search. What that does to a Sume MCP session.

If you set ENABLE_TOOL_SEARCH=false, or point Claude Code at a custom ANTHROPIC_BASE_URL, tool search is off and the MCP tool definitions from Sume's hosted server load as ordinary tools. Claude Code's MCP page lists exactly those cases, plus models earlier than the Claude 4.5 generation on Google Cloud's Agent Platform. The cost is context: Anthropic says a typical multi-server setup uses about 55k tokens of definitions before work starts, and tool search typically cuts that by over 85 percent.
What changes
With tool search on, Claude Code defers definitions and loads the ones Claude needs. With it off, everything the session can see is in the prompt on every request. For Sume the visible set depends on the credential: an API key or an OAuth session with mcp:write sees the full hosted set, and a read-only session sees fewer tools.
| Setting | Tool definitions in context | Failed servers reported |
|---|---|---|
| Default (v2.1.232+) | Only what Claude discovers | Yes |
| ENABLE_TOOL_SEARCH=false | All visible tools | No |
| Custom ANTHROPIC_BASE_URL | All visible tools | No |
| Anthropic guidance | Accuracy degrades past 30 to 50 tools | Not applicable |
Why people turn it off, and what to do instead
A gateway in front of the API may not forward the tool-search blocks, so a proxied setup is the usual reason. If that is you, trim at the source. Use a read-only OAuth session for research tasks so the write and paid tools are never listed. Use a project-scoped .mcp.json entry for the workflows that need Sume, rather than a user-wide entry on every project.
Sume's tools_list lists the tools a session can see, with safety metadata, so you can count them before choosing.
Cost angle
Larger prompts cost more on every turn. On Haiku 5.5 the line that matters is the 100,000-token prompt threshold, above which input is $0.50 per million tokens instead of $0.10. A full tool list does not reach that alone, but tool results and history can. Reread keeping the tool list small on a 1M context for the sizing habit.
None of this changes the Sume rules on writes: idempotency_key is required on write and paid tools, and max_spend_usd applies only when sent.
Measure before you decide
Count what you are paying for. Run tools_list in a session and note how many tools come back, then send a first prompt with tool search on and off and compare the input token counts your client reports. The difference is your real tool-definition overhead, which varies with the credential.
If the difference is small, leave tool search off and move on. If it is large, find a way to keep it on, such as routing through a gateway that passes the tool-search blocks, or using a narrower credential.
Sources
Related posts
More in Integrations
- Claude Code tells Claude when an MCP server fails: Sume down vs auth
With tool search on, Claude Code reports failed MCP servers to the model. How to read that when Sume's server will not connect, 401 or 403.
- Can GLM-5.3 call Sume's hosted MCP? Function calling, not MCP
Z.ai's GLM-5.3 page lists function calling and does not mention MCP. How a GLM agent reaches Sume through an MCP client or through plain HTTP.
- How to add an MCP server to ChatGPT with developer mode
Turn on ChatGPT developer mode, create an app for the server's URL, and sign in with OAuth. The steps, with Sume's hosted MCP server as the example.
- How to add subtitles to a video in Python
Add subtitles to a video in Python with Requests: POST the video URL to Sume's /v1/video-captions, poll the job, then read the captioned video_url.
Written by Sume