Plan in Claude Code over MCP, then hand the batch to bulk runs
Use Claude Code with Sume's hosted MCP to prototype one video and read costs, then run the real batch from your backend with Format bulk runs and a spend cap.

Use Claude Code with Sume's hosted MCP to find the right prompt, model and cost for one video, then move the repeatable job to your backend. Sume Format bulk runs queue up to 100 runs behind a concurrency window, and the queue runs on the server, not on your laptop. The two surfaces suit different halves of the work.
Step 1: connect and look around
Claude Code adds a remote HTTP server with claude mcp add --transport http <name> <url>, and OAuth servers sign in through /mcp or claude mcp login <name> (Claude Code docs, read 2026-10-07). For Sume, the quickstart gives the two commands:
claude mcp add --transport http sume https://mcp.sume.com/mcpclaude mcp login sume- Then ask Claude to call
tools_listormcp_healthto confirm the session.
Step 2: prototype one video
Ask the agent to read tools_schema for the tool you need, run it with dry_run=true, and report the estimate. Then submit one real call with an idempotency_key and a max_spend_usd. Keep the winning prompt, model id and settings in a note. This is cheap, interactive and easy to adjust.
Step 3: move the recipe to a Format
The hosted MCP tool list I read covers generation, jobs, assets, crawl and Avatar tools. It does not list Format runs, which are documented on the HTTP API. Turn your note into a Format recipe (style, rules and output contract), then call it from your backend.
| Stage | Surface | Why |
|---|---|---|
| Explore and price one clip | Claude Code with hosted MCP | A person reviews each step |
| Encode the look | A Format | The recipe is versioned and reused |
| Run 100 variants | POST /v1/formats/{handle}/{slug}/bulk-runs | Server-side queue with a concurrency window |
| Collect results | Poll the queue, read each child run | The queue has no webhook of its own |
Step 4: queue the batch
A bulk request holds a concurrency number and an items array, where each item is an ordinary run body. Webhooks are per item through communication.webhook_url, and queue progress comes from polling GET /v1/format-run-queues/{queue_id}. Put a generation_spend_cap_usd on every item. A queue status of completed means every item is terminal, not that every item succeeded, so read each child receipt.
Use an API key with formats:write for creation and formats:read for polling. Service-account keys cannot create Format runs or bulk queues.
Keep the agent out of the loop
Do not ask the chat agent to submit 100 paid calls one by one. The batch is a job for the queue, where cost, retries and results are all recorded.
Sources
Related posts
More in Integrations
- Video generation MCP connector: six questions before you connect
Before you let an agent spend money through any MCP connector, ask six questions. Here are Sume's hosted MCP answers on auth, scopes, idempotency and spend.
- Voiceover plus a Lyria bed for a reel in one agent session over MCP
Ask an MCP client to run tts_create and music_create on Sume, then check the take with stt_create. Which tool does which job, and where the agent must wait.
- How to add an MCP server to ChatGPT with developer mode
Turn on ChatGPT developer mode, create an app for the server's URL, and sign in with OAuth. The steps, with Sume's hosted MCP server as the example.
- How to add subtitles to a video in Python
Add subtitles to a video in Python with Requests: POST the video URL to Sume's /v1/video-captions, poll the job, then read the captioned video_url.
Written by Sume