Pydantic AI MCPToolset with Sume: which headers and caps to pass
Pydantic AI connects a remote MCP server with MCPToolset and a headers argument. Against Sume, an API key unlocks paid tools, so set caps and idempotency keys.

Connect with MCPToolset, pass your Sume key in the headers argument, and then tell the agent to attach idempotency_key and max_spend_usd to every paid call. The key matters because Sume's API-key sessions see the full tool set, paid tools included, unlike an OAuth session with only mcp:read.
What the Pydantic AI docs say
Streamable HTTP is described as the recommended way to connect to a remote MCP server. MCPToolset takes the server URL, accepts static headers such as an API key through headers, supports .prefixed('name') to avoid name collisions, and takes a custom http_client for timeouts, TLS and certificates. You register it on the agent through the toolsets parameter.
Choose the credential
Do not hard-code the key. Read it from an environment variable when you build the headers dictionary, and fail at start-up if it is empty, the same way Sume's webhook verifiers refuse an empty secret. A missing key that silently becomes a blank header produces an authentication error deep inside a run, which is harder to diagnose than a clear start-up message.
| Credential | How it is passed | What the agent sees |
|---|---|---|
| API key | headers argument; Authorization: Bearer or x-api-key | Full tool set; paid and write calls need idempotency_key |
| OAuth token | Through a client that completes MCP OAuth | mcp:read only unless Write was granted; no refresh, one-hour token |
| No credential | Not accepted | Authentication challenge |
Rules to put in the agent instructions
Use .prefixed('sume') so that tool names such as generate_image cannot clash with another server's tools, and so your logs show which server a call came from.
Rate limits apply too. Reads and submits hold separate rate budgets, and a 429 carries a retry delay that the client should honor. Give the agent a rule to wait rather than retry in a tight loop.
- Always call
dry_runbefore the first paid submit, orgeneration_admission_preview. - Always pass
max_spend_usdon a paid call, because Sume enforces it only when supplied. - Always pass a stable
idempotency_keyper intent, so a framework retry returns the same job. - Use
jobs_waitto follow a submitted job, and never resubmit while one is running.
Timeouts
Generation jobs can outlast a short HTTP timeout. The Pydantic AI docs show a custom httpx client with a timeout for this, so set one above your longest expected jobs_wait. If the agent loses the connection, the job continues on Sume's side, so resume by calling jobs_wait with the same job id rather than creating a new one.
For long scripted flows, Sume's script_run tool batches several calls in one request and reports a max_paid_calls limit you can set, which keeps a model from looping on paid calls one at a time.
Sources
More in Developers
- pytest: fail CI when a configured image model id isn't in Sume's list
A pytest that reads your image model ids and checks each against GET /v1/images/models, so a retired id such as gpt-image-1 fails CI before it fails a customer.
- Python asyncio semaphore: submit transcription jobs with a cap
Run Sume STT submits through asyncio.Semaphore and asyncio.to_thread, honor retry-after on 429, and keep one Idempotency-Key per clip. Tested code.
- Python exceptions for Sume errors: retry on the class, not the status
Map the Sume error envelope code to two exception classes, Retryable and Fatal, so one except clause drives retries. Covers 409, 429, 402 and 503.
- Python: which Sume video models take 20 seconds at 1080p?
A short Python script reads GET /v1/videos/models and filters by duration and resolution, so new launches never break a hard-coded model list.
Written by Sume