LiteLLM /mcp-rest/tools/call: test Sume tools without an LLM
LiteLLM exposes REST routes to list and call MCP tools with no model in the loop. Use them to smoke-test Sume's read tools and gates before an agent sees them.

When an agent misuses a tool, it is hard to tell whether the model, the proxy or the server is at fault. LiteLLM's MCP gateway has a way to take the model out of the question. The LiteLLM MCP page documents /mcp-rest/tools/list and /mcp-rest/tools/call, usable without an LLM, read 2026-10-03. Pointed at a Sume entry, they give you a deterministic smoke test.
What to call, in order
Start with the listing, then call only read tools. Sume's tools and gates page names mcp_health, tools_list, tools_schema, catalog_list, balance_get and usage_get among the tools. The LiteLLM page lists an x-mcp-servers header to select which servers a request targets, so send it to keep the test on the Sume entry.
| Step | Route | What you learn |
|---|---|---|
| 1 | /mcp-rest/tools/list | The proxy reaches Sume and your credential works |
| 2 | /mcp-rest/tools/call with mcp_health | The server answers a tool call |
| 3 | /mcp-rest/tools/call with balance_get | A read tool returns your own account data |
| 4 | A write tool under a read-only credential | The expected insufficient_scope error |
Test the gates on purpose
The gates are the part worth testing deliberately. Paid and write tools require an idempotency_key; calling one without it should be refused. Previews such as dry_run and generation_admission_preview show cost without spending. max_spend_usd is enforced only when you provide it, so a test that omits it proves nothing about caps.
With an OAuth credential limited to mcp:read, a write or paid tool returns insufficient_scope, per the Sume OAuth page. Seeing that error through the proxy confirms the proxy did not widen the credential.
Keep the test cheap and repeatable
Do not run a paid create in a smoke test unless you intend to pay for it. If you do, send a fixed idempotency_key so a rerun does not create a second job. Record the job id, then use jobs_wait or jobs_result to read it back; jobs_wait holds 50 seconds by default and 55 at most, and wait_slice_expired means ask again with the same ids.
LiteLLM's page mentions MCP cost tracking and guardrails but gives no detail, so this post does not describe them. Confirm any spend limit on the Sume side.
- Use the
x-mcp-serversheader so the test stays on Sume. - Call read tools first.
- Expect
insufficient_scopefor writes on a read-only credential. - Reuse one idempotency key per test case.
Sources
Related posts
More in Developers
- Live AI avatar API: a Tavus conversation vs a Sume job
A live avatar API creates a room you join. Sume's Avatar API creates a job you poll. Field-by-field map of Tavus create conversation and Sume talking-video.
- LM Studio mcp.json: add Sume's hosted MCP server with an API key
LM Studio 0.3.17 and later accept remote MCP servers in mcp.json. The exact entry for https://mcp.sume.com/mcp with a Bearer header, plus the gates to know.
- Load-test your Sume webhook receiver with a signed burst
Fire 500 correctly signed job.completed requests at your own receiver before a bulk run does it for real. A runnable Python harness and the 10-second budget.
- LRC lyrics file from an AI music track: Sume STT segments in Python
YouTube accepts .lrc lyric files. A short Python script turns STT sentence segments from a generated song into [mm:ss.xx] lines. Check results before upload.
Written by Sume