Semantic Kernel MCPStreamableHttpPlugin with Sume: load_prompts False
Connect Semantic Kernel to Sume's hosted MCP with MCPStreamableHttpPlugin. Set load_prompts to False because Sume serves only tools, and size the timeout.

To use Sume from Semantic Kernel, add Sume's hosted MCP endpoint (https://mcp.sume.com/mcp) with MCPStreamableHttpPlugin and set load_prompts to False. Microsoft's documentation, read on 2026-10-04, warns that prompt loading can hang when a server has none, and Sume's server exposes tools only. Then raise request_timeout above what your longest tool call needs, because one jobs_wait call on Sume can hold for 55 seconds.
What does Semantic Kernel offer?
The Semantic Kernel MCP plugins page lists three plugin types: MCPStdioPlugin, MCPSsePlugin and MCPStreamableHttpPlugin. You install the extra with pip install semantic-kernel[mcp]. The options are load_tools, which defaults to true, load_prompts, which also defaults to true, and request_timeout, for which the page does not state a unit.
Which settings suit Sume?
| Option | Setting for Sume | Why |
|---|---|---|
| Plugin type | MCPStreamableHttpPlugin | the hosted endpoint is a remote HTTPS URL |
load_tools | True | tools are all Sume exposes |
load_prompts | False | Sume serves no prompts and the docs warn of a hang |
request_timeout | above 55 seconds, check the unit | jobs_wait holds up to 55 seconds per slice |
How should I authenticate?
For an unattended agent, use an API key, sent as Authorization: Bearer <SUME_API_KEY> or in an x-api-key header, as MCP OAuth and API keys describes. I have not verified how the plugin takes headers, so check Microsoft's page for the parameter before wiring it in. An API-key session sees the full tool set, so keep the number of functions the kernel can call small and require idempotency_key on every write and paid call.
OAuth suits an interactive app. A mcp:read grant hides write and paid tools, which is the right default for a copilot that only inspects jobs and usage.
What is the first call?
Start with mcp_health, which reports endpoint readiness, auth source and safety posture, then tools_list. For any generation, call the create tool and then jobs_wait and jobs_result, never a wait loop inside the create. The MCP quickstart walks through those calls, and Jobs and results covers statuses.
Plan the timeout for the slowest call you make, not the average. A wait slice that returns with jobs unfinished is normal, so call it again with the ids that remain rather than raising the timeout without limit.
How do I keep the agent safe?
Limit what the kernel may call. Sume's paid tools need an idempotency_key, and max_spend_usd caps a call only when it is sent, so have your own code, not the model, set the cap and build the key from the work. Start with read tools such as catalog_list and usage_get, then add one create tool once you have seen a full create, wait and result cycle succeed. That order keeps the first mistakes cheap.
Sources
Related posts
More in Integrations
- Shopify selling_plan_id in order webhooks: a Sume welcome clip
Shopify 10.01 order webhooks carry selling_plan_id on line items. Use it to start one Sume welcome video per subscription order, with a safe retry key.
- A Supabase job table for Sume webhooks: upsert on job_id
Sume retries webhooks up to 10 times. A Postgres table keyed on job_id with ON CONFLICT turns repeat deliveries into no-ops. Schema, SQL and the order of steps.
- TikTok Ads MCP: 400-tool or 40-tool URL beside Sume's hosted MCP?
TikTok's Agentic Hub has a Full Disclosure URL (about 400 tools) and a Progressive one (about 40). Which to pick with Sume's hosted MCP server also connected.
- ubuntu-latest moves to 26.04: test your Sume workflow now
GitHub is moving ubuntu-latest to Ubuntu 26.04 between Oct 19 and Nov 19. Run a Sume API smoke test on both images, or pin 24.04, before it flips.
Written by Sume