Your agent calls generate_video, or Sume's agent does: two meters

With hosted MCP your model client runs the loop and you pay its provider; with an Agent Completion Sume's agent runs it and you set generation_spend_cap_usd.

5 min readSume
All posts

There are two ways an AI agent can produce video through Sume, and they put the loop in different places. With hosted MCP, your own client and model, say Claude Haiku 5.5, drive the loop and call tools like generate_video. With an Agent Completion, Sume's agent runs the loop and you only send the task and a generation_spend_cap_usd. Either way generation is charged on Sume's rates; the difference is who chooses the model, who holds the tool loop, and where the cap lives.

The docs say an Agent Completion accepts only model: sume-agent, so you cannot select Haiku 5.5, Mistral Large 4 or any other model for it. If you need a specific model reasoning over the tools, use the MCP route with a client that runs that model.

Side by side

From the MCP, Agent Completions and Safe automation pages, read 2026-10-09.

Two ways to drive generation from an agent
QuestionHosted MCP from your clientAgent Completion
Who runs the model loopYour client and chosen modelSume's agent (sume-agent)
Entry pointhttps://mcp.sume.com/mcpPOST /v1/agent/completions
AuthOAuth or API keyAPI key with agent_completions:write
Spend controlOptional max_spend_usd, dry_run, wallet admissionRequired generation_spend_cap_usd
ResultTool results in your clientAsync receipt; poll status_url or use a webhook
Needs a person presentInteractive OAuth, or a key for automationNo person has to monitor it

Choosing

Pick hosted MCP when the model matters to you: a cheap small model for routing, a stronger one for planning, or a client your team already uses such as Cursor or Claude Code. Pick an Agent Completion when your backend only has a task and you want Sume to carry out the steps with a hard cap on the generation spend.

Neither is the Studio Agent product, which is not a public MCP connector. Do not point a client at it.

What each meter covers

The MCP route has two meters. The model provider bills tokens at its own prices, for example Haiku 5.5 at $0.10 input and $0.50 output per million tokens for short prompts, and Sume bills generation at its catalog rates. The Agent Completion's documented cap covers generation spend, and the page does not describe a separate token charge, so check the pricing page before you assume anything about it. For a cap, start from the catalog arithmetic: three 5-second Wan 3.0 720p clips are 3 x 5 x $0.125 = $1.875, so a cap of $2.00 covers them with a small margin.

A checklist for either path

Whichever route you take, treat spend as something you configure, not something you hope for.

  • MCP: require idempotency_key and max_spend_usd in the instructions, and preview with dry_run before a burst.
  • Agent Completion: pick the cap from the catalog price of the clips you expect, plus a margin, and send Idempotency-Key on the request.
  • Both: read the result from durable media.sume.com URLs, and keep signed URLs and keys out of logs.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume