Your agent calls generate_video, or Sume's agent does: two meters
With hosted MCP your model client runs the loop and you pay its provider; with an Agent Completion Sume's agent runs it and you set generation_spend_cap_usd.

There are two ways an AI agent can produce video through Sume, and they put the loop in different places. With hosted MCP, your own client and model, say Claude Haiku 5.5, drive the loop and call tools like generate_video. With an Agent Completion, Sume's agent runs the loop and you only send the task and a generation_spend_cap_usd. Either way generation is charged on Sume's rates; the difference is who chooses the model, who holds the tool loop, and where the cap lives.
The docs say an Agent Completion accepts only model: sume-agent, so you cannot select Haiku 5.5, Mistral Large 4 or any other model for it. If you need a specific model reasoning over the tools, use the MCP route with a client that runs that model.
Side by side
From the MCP, Agent Completions and Safe automation pages, read 2026-10-09.
| Question | Hosted MCP from your client | Agent Completion |
|---|---|---|
| Who runs the model loop | Your client and chosen model | Sume's agent (sume-agent) |
| Entry point | https://mcp.sume.com/mcp | POST /v1/agent/completions |
| Auth | OAuth or API key | API key with agent_completions:write |
| Spend control | Optional max_spend_usd, dry_run, wallet admission | Required generation_spend_cap_usd |
| Result | Tool results in your client | Async receipt; poll status_url or use a webhook |
| Needs a person present | Interactive OAuth, or a key for automation | No person has to monitor it |
Choosing
Pick hosted MCP when the model matters to you: a cheap small model for routing, a stronger one for planning, or a client your team already uses such as Cursor or Claude Code. Pick an Agent Completion when your backend only has a task and you want Sume to carry out the steps with a hard cap on the generation spend.
Neither is the Studio Agent product, which is not a public MCP connector. Do not point a client at it.
What each meter covers
The MCP route has two meters. The model provider bills tokens at its own prices, for example Haiku 5.5 at $0.10 input and $0.50 output per million tokens for short prompts, and Sume bills generation at its catalog rates. The Agent Completion's documented cap covers generation spend, and the page does not describe a separate token charge, so check the pricing page before you assume anything about it. For a cap, start from the catalog arithmetic: three 5-second Wan 3.0 720p clips are 3 x 5 x $0.125 = $1.875, so a cap of $2.00 covers them with a small margin.
A checklist for either path
Whichever route you take, treat spend as something you configure, not something you hope for.
- MCP: require
idempotency_keyandmax_spend_usdin the instructions, and preview withdry_runbefore a burst. - Agent Completion: pick the cap from the catalog price of the clips you expect, plus a margin, and send
Idempotency-Keyon the request. - Both: read the result from durable
media.sume.comURLs, and keep signed URLs and keys out of logs.
Sources
Related posts
More in Agents
- Planlock in front of Sume MCP: approve the plan, then the calls
Planlock is an MCP proxy that enforces a human-approved plan. Where it fits in front of Sume's hosted MCP, and which Sume gates still do work behind it.
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
- Safe automation for AI agents that call paid APIs
Keep agents read-only by default, keep secrets out of logs, and on hosted MCP send an idempotency_key, preview with dry_run, and cap with max_spend_usd.
- Scheduled AI video agent runs: cron, API triggers, and receipts
A Sume schedule is a saved Agents automation that runs on a cron cadence and returns a run receipt. Author it in the dashboard; start and monitor runs by API.
Written by Sume