AutoGen MCP workbench: connect McpWorkbench to Sume
Use AutoGen's McpWorkbench with Sume's hosted MCP server: StreamableHttpServerParams, an API-key header, a 60-second timeout, and fewer tools.

McpWorkbench is AutoGen's wrapper around one MCP server: it lists and calls the server's tools, and you give it to an AssistantAgent as its workbench. For a remote server such as Sume's hosted MCP, build StreamableHttpServerParams with url="https://mcp.sume.com/mcp", your Sume API key in an Authorization: Bearer header, and timeout raised from its default of 30 seconds to 60, because one Sume jobs_wait call can hold for 55.
AutoGen's side comes from its autogen_ext.tools.mcp API reference; Sume's side comes from MCP OAuth and API keys, MCP tools and gates, and Jobs and results, all read on 2026-09-28. Sume has no AutoGen package or extension: this is AutoGen's own MCP client talking to Sume's remote server, and Sume's basics page says hosted MCP still works but is not part of the primary path today. Pydantic AI MCP server covers another Python framework.
How do I connect McpWorkbench to Sume's MCP server?
Install the MCP extra, autogen-ext[mcp], alongside the AutoGen packages the example imports. AutoGen says to use the workbench as a context manager, so its MCP session is initialized and cleaned up properly; start() and stop() do the same by hand. Read the key from the environment, and have the agent call mcp_health and tools_list, Sume's read-only discovery tools, before anything else:
import asyncio
import os
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient
from autogen_ext.tools.mcp import McpWorkbench, StreamableHttpServerParams
async def main() -> None:
params = StreamableHttpServerParams(
url="https://mcp.sume.com/mcp",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=60.0,
sse_read_timeout=300.0,
)
async with McpWorkbench(params) as sume:
agent = AssistantAgent(
"media_assistant",
model_client=OpenAIChatCompletionClient(model="gpt-4.1-nano"),
workbench=sume,
)
await Console(agent.run_stream(task="Call mcp_health and tools_list, then summarize."))
asyncio.run(main())What do timeout and sse_read_timeout control?
AutoGen's example labels timeout the HTTP timeout in seconds, default 30.0, and sse_read_timeout the SSE read timeout in seconds, default 300.0, or 5 minutes. The reference doesn't say which one bounds a long tool call, so keep both above the 55 seconds a Sume jobs_wait can hold: raise timeout to 60 and leave sse_read_timeout at 300.
A client-side timeout does not cancel a Sume job; it keeps running and billing. For a render longer than one wait, the agent should call jobs_wait again with the same ids on wait_slice_expired and never resubmit the paid create. MCP tool call timeouts on long-running video jobs has the pattern.
| `StreamableHttpServerParams` field | AutoGen default | For Sume |
|---|---|---|
url | Required | https://mcp.sume.com/mcp |
headers | None | Authorization: Bearer <key> or x-api-key |
timeout | 30.0 | 60.0, above the 55-second jobs_wait hold |
sse_read_timeout | 300.0 | Leave at 300.0 |
terminate_on_close | True | Leave as is |
How do I give the agent only some Sume tools?
The workbench offers the agent every tool the server lists, and its reference documents no filter, only tool_overrides for a tool's name and description. An API-key session lists Sume's full hosted tool set, paid tools included, and AutoGen's MCP reference describes no approval step before a call. To narrow it, call mcp_server_tools() with the same parameters instead: it returns one adapter per tool, adapters go straight into an agent's tools list, and each has a name to keep or drop.
Sume's docs say to give agents read-only commands first and require explicit confirmation before paid generation, so start with read tools and add generate_image only when unattended spend is acceptable. Each paid call needs an idempotency_key; dry_run=true previews admission and cost without submitting; max_spend_usd caps a call only when it is sent. In the example above, this replaces the async with block:
# at the top of the file
from autogen_ext.tools.mcp import mcp_server_tools
# in main(), instead of the async with block
# add "generate_image" only when unattended spend is acceptable
ALLOWED = {"tools_schema", "jobs_status", "jobs_wait", "jobs_result"}
tools = [t for t in await mcp_server_tools(params) if t.name in ALLOWED]
agent = AssistantAgent(
"media_assistant",
model_client=OpenAIChatCompletionClient(model="gpt-4.1-nano"),
tools=tools,
)
await Console(agent.run_stream(task="Call tools_schema for generate_image."))Can I use the workbench's resources and prompts with Sume?
No. McpWorkbench supports tools (list_tools, call_tool), resources, resource templates, and prompts, plus optional sampling through model_client. Sume's current server declares only the tools capability and answers other request methods with JSON-RPC error -32601 and a message that starts MCP method is not supported. So list_resources, read_resource, list_prompts, and get_prompt fail against Sume; stick to list_tools and call_tool. MCP tools vs resources vs prompts explains the three.
Sources
Related posts
More in Integrations
- AWS Lambda maximum concurrency for SQS and AI API jobs
Lambda's SQS maximum concurrency caps how many instances one queue can invoke, from 2 to 1,000. How to set it so AI jobs stay within API limits.
- Azure AI Foundry MCP tool: connect an agent to Sume
Add an MCP tool to an Azure AI Foundry agent: a Custom keys connection for Sume's API key, then server_url, allowed_tools, and require_approval.
- Azure DevOps scheduled pipeline: a nightly cron in UTC
Add a schedules block with a UTC cron to the pipeline YAML, set always: true to run without code changes, and key any paid API call to the date.
- BullMQ retry: exponential backoff for paid API jobs
Set attempts and an exponential backoff on BullMQ jobs, stop early with UnrecoverableError, and key each paid API call to the job so retries replay.
Written by Sume