Copilot CLI with a local Ollama model: do Sume MCP tools still work?
Copilot CLI can now discover Ollama models, but a local model is not offline mode. What that means for Sume's remote MCP tools, spend gates and tool calling.

A local model in GitHub Copilot CLI can drive Sume's hosted MCP tools, provided the model supports tool calling and streaming, and the tools still run on Sume's servers, not on your machine. GitHub's changelog entry for Copilot CLI 1.0.94 and later says local models are discovered through the /model command, and that selecting one does not turn on an offline-only mode.
The GitHub facts come from the changelog post Discover local models in GitHub Copilot CLI, read on 2026-10-11. Sume's facts come from the MCP overview, tools and gates and OAuth and API keys pages. The post says nothing about MCP servers, so the combination below is my reading, and the steps that depend on it are marked as things to test.
What does the GitHub post actually say?
The changelog says the /model command can now list models from running Ollama instances alongside configured and cloud models. Discovery does not add anything automatically: you review the provider details and endpoint, then choose "Add and use for this session" or "Add without switching". The models must already be installed, since the CLI does not install runtimes or download models. Ollama is the only local runtime the post names.
Two sentences matter for Sume. First, "Models must support tool calling and streaming." Second, GitHub says choosing a local model does not mean offline: GitHub telemetry continues, remote providers can still receive prompts over the network, and offline mode needs COPILOT_OFFLINE=true.
Where does each part of the call run?
Sume's tools are server-side. A tool call leaves your session as an HTTPS request to https://mcp.sume.com/mcp, and generation, job state and the wallet all live on Sume. The local model only decides which tool to call and with what arguments. Swapping a cloud model for a local one therefore changes the quality of those decisions, not what the tools do or what they cost.
| Part | Where it runs | What the source says |
|---|---|---|
| Choosing the tool and arguments | Local Ollama model | GitHub: model must support tool calling and streaming |
| Prompts | Local, but remote providers can still receive them | GitHub: not offline mode |
| Sume tool execution | Sume, at mcp.sume.com | Sume MCP overview |
| Generation spend | Sume wallet and admission | Sume tools and gates |
| Fully offline use | Needs COPILOT_OFFLINE=true | GitHub; effect on remote MCP not stated |
What happens with COPILOT_OFFLINE=true?
The GitHub post names the variable but does not say what it does to remote MCP servers. Do not assume either way. A remote streamable HTTP server cannot work without a network, so if the setting blocks outbound calls, Sume's tools will fail with a connection error, and if it does not, they will run. Test with mcp_health, which Sume lists as a read-only readiness check, and take the answer from what you see.
Either result is safe for your wallet. A tool call that never reaches Sume cannot create a job, and a repeated call uses the same idempotency_key so it cannot create two.
How do I keep a small local model from overspending?
Smaller models make more argument mistakes, so use Sume's own gates rather than the model's judgement. Connect with OAuth mcp:read and the model sees only read-only tools; a write or paid call returns insufficient_scope. When you need generation, add mcp:write and require dry_run=true first, then submit with max_spend_usd. Sume enforces that cap only when you send it, so put it in your prompt or wrapper, not in the hope the model remembers.
Sume's docs also recommend generation_admission_preview before expensive bursts. A local model that loops on jobs_wait is harmless because the call is read-only and holds at most 55 seconds, but the same loop on generate_video is not, which is why the idempotency key is required on every paid call.
Sources
Related posts
More in Integrations
- gemini mcp add for Sume: transport, header, timeout, include-tools
One gemini mcp add command registers Sume's hosted MCP over HTTP with a key header, a timeout above 55 seconds, and an include-tools list for read-only use.
- Jira Rovo MCP plus Sume hosted MCP: turn a ticket into a release clip
Use Atlassian Rovo MCP and Sume hosted MCP in one session: read a Jira ticket, dry-run a clip, cap spend, and post the media.sume.com link back.
- Mastra Connect 1.0: no MCP approval by default, Sume not listed
Mastra Connect 1.0 stopped asking approval for discovered MCP tools, and its docs list seven MCP providers, not Sume. Use MCPClient and gate paid tools.
- n8n MCP Client Tool: Tools to Include modes for Sume
n8n's MCP Client Tool can expose All, Selected, or All Except tools. How to pick the set for Sume so an AI Agent node reads freely but cannot spend by accident.
Written by Sume