Sort your MCP tool allowlist: stable order keeps the prompt cache warm
MCP 2026-07-28 asks servers for ttlMs, cacheScope and a deterministic tool order. On the client, sort your media tool allowlist the same way. Python snippet.

Sort the tool list you hand to the model, and never let its order depend on when a server happened to answer. MCP 2026-07-28 now requires tools/list, prompts/list, resources/list and resources/read to return ttlMs and cacheScope (public or private), and says servers SHOULD return tools in a deterministic order so prompt caches hit. A client that builds its own allowlist should do the same. Source: the changelog, read 2026-10-03.
The three cache fields
Prompt caches match on a prefix. Tool definitions sit near the start of a prompt, so a reordered tool list invalidates everything after it. The spec changes target exactly that.
| Field or rule | Meaning |
|---|---|
| ttlMs | How long a list result may be reused |
| cacheScope: public | Result may be shared |
| cacheScope: private | Result is specific to the caller |
| Deterministic order (SHOULD) | Same tools, same order, so a cached prompt prefix still matches |
Why a media server needs the scope flag
A Sume session does not always see the same tools. The OAuth docs say an mcp:read session sees only read-only tools, a session with mcp:write sees mutating and paid tools too, and an API key sees the full hosted set. A list taken under one grant is wrong under another, which is the case a private scope describes. The docs also tell you to call tools_list for the session-visible subset instead of assuming it.
So cache per session and grant, and do not share a list between a read-only connector and a write-enabled one.
Pin and sort on the client
If you wrap hosted tools in your own agent, choose the few you need and sort them by name before building the prompt. A short allowlist also keeps the prompt small. The tool ids below are live ids from the Sume tools page; check tools_list for your session before relying on any of them.
ALLOW = {"generate_image", "generate_video", "jobs_wait",
"jobs_result", "tools_schema"}
def pinned(tools: list[dict]) -> list[dict]:
picked = [t for t in tools if t["name"] in ALLOW]
return sorted(picked, key=lambda t: t["name"])
demo = [{"name": "jobs_wait"}, {"name": "account_me"}, {"name": "generate_image"}]
print([t["name"] for t in pinned(demo)])Watch for churn
Sorting fixes order. It does not fix content. If a tool description changes, the prefix changes with it, so avoid putting timestamps, balances or per-request values in tool descriptions. Keep dynamic facts, such as the current balance, in a tool result where they belong.
Sources
Related posts
More in Developers
- Split a narration script by model character limit: Python
ElevenLabs lists a 40,000 character limit for Flash v2.5 and 10,000 for v4 and v4 Turbo. A Python splitter that cuts on sentences, with a concat step on Sume.
- Patching Supabase Postgres 17.11 vs Sume's 10-attempt webhook budget
Supabase's September 25 Postgres 15.19 and 17.11 releases fix 44 CVEs. A restart can outlast Sume's ten 30-second webhook attempts, so plan a redeliver.
- Supabase cached egress is $0.03/GB: cost of serving a 20 MB AI clip
Supabase lists cached Storage egress at $0.03 per GB. Worked arithmetic for serving generated clips, and when to link a Sume media URL instead of copying.
- End-user id on jobs: OpenAI safety identifier vs Sume metadata
OpenAI's Realtime guide asks for an OpenAI-Safety-Identifier header. Sume stores caller metadata on the job but does not send it to the provider. Use both.
Written by Sume