smolagents MCPClient: give a CodeAgent Sume's remote tools
Connect a smolagents CodeAgent to Sume's hosted MCP with MCPClient and the streamable-http transport, send an API key header, and wait on jobs correctly.

Pass MCPClient a dict with "url": "https://mcp.sume.com/mcp" and "transport": "streamable-http", use it as a context manager, and hand the tools it yields to a CodeAgent. That is the Streamable HTTP form the smolagents tools tutorial shows for a local server, pointed at Sume's hosted endpoint instead.
The part the tutorial does not cover is authentication and paid calls. This post fills that in from Sume's side.
What does the smolagents page say about remote servers?
From the tutorial, read on 2026-10-02: MCPClient loads tools from an MCP server and gives control over the connection. For Streamable HTTP servers you pass a dict, and for ToolCollection.from_mcp the page says the dict holds parameters for mcp.client.streamable_http.streamablehttp_client plus a transport key set to "streamable-http". It also warns to verify the source of any MCP server, since stdio servers run code on your machine and even remote ones deserve caution.
Sume's hosted server is the remote case, so nothing runs locally except your agent. Because the dict feeds the Streamable HTTP client, a headers entry is the natural place for a credential; confirm that against the mcp package version you have installed.
How do you send the Sume API key?
Sume accepts Authorization: Bearer <key> or an x-api-key header on the hosted endpoint (OAuth and API keys). OAuth is the interactive path and needs a browser consent step, so a script uses a key. An API-key session sees the full hosted tool set, paid tools included. Create the key in the dashboard and read it from the environment; never put it in the prompt.
import os
from smolagents import CodeAgent, InferenceClientModel, MCPClient
params = {
"url": "https://mcp.sume.com/mcp",
"transport": "streamable-http",
"headers": {"x-api-key": os.environ["SUME_API_KEY"]},
}
with MCPClient(params) as tools:
agent = CodeAgent(tools=tools, model=InferenceClientModel())
print(agent.run("Call mcp_health and report the auth source and safety posture."))What does a CodeAgent do differently with these tools?
A CodeAgent writes Python that calls tools as functions, so it can loop. That is useful and risky with a paid server. A loop that calls a create tool once per scene is the shape Sume's own script_run exists for, but inside smolagents the model writes the loop and nothing else caps it. Sume's docs require an idempotency_key on every paid or write tool, so tell the agent to build a distinct key per call from the loop index, for example "scene-" + str(i).
Keep the tool list small. The tutorial itself warns that too many tools can overwhelm weaker models. Sume's hosted set is broad (generation, crawl, avatars and more), so after the first tools_list check, pass the agent only the families the task needs.
How should the agent handle slow jobs?
Generation returns job ids, not files. The agent should call jobs_wait with the ids and then jobs_result. One wait holds for 55 seconds at most (default 50) and may answer wait_slice_expired; the correct reaction is to call jobs_wait again with the same ids. Never rerun the create, which would be a second paid job unless the same idempotency_key is replayed (Jobs and results).
Batch the wait: job_ids takes 1 to 20 ids, so one call covers a whole fan-out. Put that rule in the agent's instructions, since a code-writing model otherwise tends to poll each id in its own loop.
- Preview first:
generation_admission_previewordry_run=truereturns an estimate without submitting. - Cap it:
max_spend_usdis enforced only when the agent passes it. - Report media URLs, not signed links or the key, in the final answer.
Should you turn on structured_output?
The tutorial describes a structured_output=True option on MCPClient that enables outputSchema support, so the agent's model can see the shape of a tool's result before calling it. It defaults to False today, with a note that the default will change to True in a future release, and the page recommends setting it explicitly.
Whether it helps with Sume depends on whether a given Sume tool declares an output schema. Check the contract that tools_schema returns for the tool you plan to use before you rely on it; without a schema you simply get text back, as before. Either way, pass the flag explicitly so a library upgrade does not change how your agent reads results.
Finally, keep the context manager. The tutorial shows a manual try/finally with disconnect() for cases where you cannot use with, and the same rule applies: close the client when the agent finishes. Closing the client does not cancel anything on Sume's side. A generation job you already submitted keeps running and billing, so read its jobs_status before you walk away.
Sources
Related posts
More in Integrations
- Snap Marketing API ai_content_source: the AI declaration on Media
Snap added an ai_content_source attribute to the Media entity in its Marketing API on August 28, 2026. What the changelog says, and how to record it.
- Supabase Queues visibility timeout as a Sume job poller
Supabase Queues (pgmq) hide a read message for a visibility timeout, which fits polling a Sume job: read, check status, archive on terminal, else it reappears.
- TikTok max_video_post_duration_sec: trim before Direct Post
TikTok's creator_info endpoint returns max_video_post_duration_sec, a per-creator limit. Read it first, then cut the clip to fit with Sume video trim.
- TikTok privacy_level_options: public vs private account values
TikTok creator_info returns different privacy_level_options for public and private accounts, and Direct Post privacy_level must match one. Values for each.
Written by Sume