Agents SDK cache_tools_list: stale tools after you grant Sume Write
With cache_tools_list on, an OpenAI Agents SDK MCP server can keep an old tool list. After a Sume Write grant, call invalidate_tools_cache() to see paid tools.

Yes, a cached list can hide tools you just unlocked. The OpenAI Agents SDK docs say every MCP server class exposes cache_tools_list, and that to force a fresh list you call invalidate_tools_cache() on the server instance. Sume's visible tools depend on the token scope, so do that after you grant Write.
Both statements come from the SDK docs (read 2026-10-07) and Sume's docs. The behavior of your exact SDK version is worth one test run.
Why Sume's list can change
Under OAuth with mcp:read only, Sume shows read-only tools. Tools that change data and paid tools appear once the session has mcp:write. With an API key you see the full set from the start, so the cache cannot go stale in that direction.
The risky case is a long-lived process that connected with a read token, cached the list, and was later given a write token by a human.
What the SDK docs say
The page states that remote servers can add noticeable latency, so all MCP server classes expose cache_tools_list, and it advises enabling the cache only if you are confident that tool definitions do not change often. It also notes that list_tools() returns detached copies of cached definitions.
For a scope-dependent server, the definitions do change with the grant, so the advice points to leaving the cache off or invalidating it deliberately.
Code to test it
The script lists tools with the cache on, invalidates it, and lists again. It uses an API key from the environment, so both counts match; the point is the call order. With an OAuth bearer token you would see the second count rise after a Write grant.
import asyncio, os
from agents.mcp import MCPServerStreamableHttp
async def main():
key = os.environ["SUME_API_KEY"]
params = {
"url": "https://mcp.sume.com/mcp",
"headers": {"Authorization": f"Bearer {key}"},
}
async with MCPServerStreamableHttp(
name="sume", params=params, cache_tools_list=True
) as server:
first = await server.list_tools()
server.invalidate_tools_cache()
second = await server.list_tools()
print(len(first), len(second))
asyncio.run(main())A rule that avoids the problem
Choose by credential. For an unattended agent with an API key, caching is fine because the set is fixed. For an interactive agent on OAuth, either leave caching off or invalidate on any reconnect.
| Credential | Tool set | cache_tools_list |
|---|---|---|
| API key | Full, fixed | Safe to enable |
| OAuth mcp:read | Read-only | Enable only if the grant never changes |
| OAuth mcp:read + mcp:write | Full | Invalidate after any re-consent |
What a stale list looks like
The model will say it has no tool for the task, or it will pick a read tool that cannot do the job. Nothing errors, since the tool list is simply short. That makes it a quiet failure. If an agent says it cannot generate after you granted Write, reconnect or invalidate the cache before debugging the prompt.
The opposite also exists. If you revoke Write, a cached list may still offer generate_video. The call then returns insufficient_scope from Sume, which is the real gate.
Guard the paid calls anyway
A fresh list does not make a paid call safe. Keep idempotency_key on each create, run dry_run first and set max_spend_usd in the arguments. The SDK also offers hosted MCP approval modes such as require_approval, which you can use to make a person confirm each paid call.
Sources
Related posts
More in Developers
- X-OpenRouter-Idempotency-Key vs Sume webhook dedupe on job_id
OpenRouter's webhook dedupe key is job_id plus status. Sume says dedupe on job_id. A SQLite sample builds a job_id plus event key that skips replays.
- Pandas DataFrame to a Sume bulk queue in 100-row chunks
Turn a product DataFrame into Sume Format bulk queues: one item per row, 100 rows per queue, a stable key per chunk, and SKU order saved beside each queue id.
- Pin the model id in an ad test: sume/auto follows the catalog
sume/auto is a pure function of the request plus the catalog version, so two ad arms made weeks apart can land on different models. Pin an explicit id in tests.
- Poll hundreds of AI jobs without a thundering herd: jitter and budgets
Poll many Sume jobs without synchronized bursts: jitter, next_poll_after_seconds, per-plan read budgets, and the math on how much polling a plan can absorb.
Written by Sume