Letta MCP server: register Sume with mcp_servers.create
Register Sume's hosted MCP in Letta with mcp_servers.create, the streamable_http type and an Authorization token, then attach its tools to an agent by id.

Register Sume in Letta with client.mcp_servers.create, set mcp_server_type to streamable_http, server_url to https://mcp.sume.com/mcp, and pass the credential through auth_header and auth_token. Then attach the tools you want to an agent by id. Letta's own guide recommends Streamable HTTP for remote servers because it supports auth headers.
What follows covers the registration call, how Letta's execution model changes what an agent can see, and the Sume details that matter once a tool call has to wait for media.
How does Letta register a remote MCP server?
From Letta's MCP tools guide, read on 2026-10-02: registration needs a server URL, optional credentials and a server type. The transport table lists Streamable HTTP as supported on the Letta API and Docker and recommended, SSE as legacy compatibility only, and stdio as local development only. Authentication can be a bearer token in auth_token, extra custom_headers, or agent-scoped variables written like Bearer {{USER_API_KEY | default_key}}.
import os
from letta_client import Letta
client = Letta()
server = client.mcp_servers.create(
server_name="sume",
config={
"mcp_server_type": "streamable_http",
"server_url": "https://mcp.sume.com/mcp",
"auth_header": "Authorization",
"auth_token": f"Bearer {os.environ['SUME_API_KEY']}",
},
)
print(server)Which Sume credential goes in auth_token?
Use a Sume API key. Sume's hosted MCP also supports OAuth, but that flow needs a person to sign in on https://mcp.sume.com/oauth/consent, and auth_token is a static string, so there is no step where a user could do that. The key goes in as Bearer <key>; Sume accepts the Authorization: Bearer form and an x-api-key header (OAuth and API keys).
The agent-scoped variable form is interesting for multi-user products: Letta's page shows a placeholder in the token string, so each agent can carry its own key. Sume keys are workspace credentials and carry wallet spend, so give each customer's agent a key tied to that customer's own Sume account, not a shared one.
Where do the tool calls run, and what does the agent see?
Letta's guide says MCP tools execute on the external server, not in Letta's sandbox: the agent sends a tool call to the Letta server, Letta forwards it to Sume, and the agent receives the result without ever holding the credential. That is a good property for a paid server, because the model cannot print or leak the key.
You attach tools to an agent with tool_ids, so an agent only has the Sume tools you chose. Start from the tools_list output, which includes safety metadata, and attach a small set. A reasonable first agent for media gets generation_admission_preview, one create such as generate_image, jobs_wait and jobs_result. Everything else on the hosted catalog, from crawl tools to timeline composition, can stay unattached until a task needs it (MCP tools and gates).
What about slow generations and retries?
Paid and write tools require an idempotency_key. If Letta or your network retries a call after a timeout, the same key replays the original receipt instead of creating a second billable job. Tell the agent to derive keys from the business intent, such as the post and revision, and not from a random value per attempt.
Media is not instant. jobs_wait holds for at most 55 seconds per call and answers wait_slice_expired at the end of a slice; call it again with the same ids (Jobs and results). The Letta guide read for this post does not state a per-call timeout for MCP tools, so check your deployment's setting and keep it above that 55-second slice. max_spend_usd only applies when the agent passes it, so put a cap in the agent's system prompt.
How do you scope what each Letta agent can spend?
The registration call is global to the Letta server, but the tools are per agent. That split is useful: register Sume once, then build agents with different tool_ids for different jobs. A research agent gets crawl_scrape, crawl_search and jobs_status. A creative agent gets the admission preview, one create and the job reads. An operations agent gets balance_get and usage_get and nothing paid at all.
Because the key is workspace-wide on Sume's side, the narrowing happens in Letta, not in Sume. If you need Sume itself to refuse writes, connect with OAuth read-only (mcp:read) instead, where a paid call returns insufficient_scope; that works only if your Letta deployment can complete the browser consent, which the page read for this post does not describe. Re-run tools_list after any change so the attached ids still match what the session can see.
Sources
Related posts
More in Integrations
- LinkedIn Videos API templateName and linkbackContext (202602+)
LinkedIn added templateName and linkbackContext to initializeUpload from version 202602. They credit the tool that made a video. What to send for a Sume render.
- Mistral connectors: confirm Sume's paid tools before they run
Add Sume as a Mistral custom MCP connector, then use tool_configuration include and requires_confirmation to keep paid generation behind a human check.
- n8n Data table as a Sume job ledger: the 200 MiB default limit
Store each Sume job id in an n8n Data table and upsert it when the webhook arrives. The default cap is 200 MiB per instance, and a full table errors inserts.
- n8n Execution Data node: find a run by its Sume job id
Save the Sume job id with n8n's Execution Data node and you can search the Executions list by it. Keys cap at 50 characters, values at 512, plan limits apply.
Written by Sume