MCP agent: trim, then captions? video-captions create is not listed

On Sume's hosted MCP, video_trim is a listed tool but captions are not creatable from tools_list. Use the REST endpoint for the burn step in an agent chain.

4 min readSume
All posts

An MCP agent on Sume's hosted server can generate and trim a clip, but it cannot rely on a captions create tool. The tool inventory lists video_trim, timeline_create and video-captions_get, and it says the legacy video-captions_create and video-caption-overlay_create are unlisted: in-flight clients can still call them by name, but they are not in tools_list. For a new agent, burn captions with POST /v1/video-captions from your own code, or fold the step into a workflow that your code controls.

What the hosted MCP does list for this chain

Paid generation tools such as generate_video need idempotency_key, and the docs say to omit payload.model to route to sume/auto unless the user named a model family. sume/auto creates 3 to 10 second clips by default, so a pinned model is the route to a 30-second render. video_trim is a write that needs idempotency_key and, under OAuth, mcp:write. The flow is video_trim, then jobs_wait, then jobs_result.

Chain steps against the hosted MCP tool list (read 2026-10-05)
StepHosted MCP toolListed in tools_list?
Generategenerate_videoYes, paid, needs idempotency_key
Waitjobs_waitYes, read
Trimvideo_trimYes, write, needs idempotency_key
Read the cutjobs_resultYes, read
Place on a timelinetimeline_createYes
Burn captionsvideo-captions_get onlyCreate is legacy and unlisted

A pattern that works

Let the agent do generate, wait and trim over MCP, and have it return the trimmed video_url. Your own service then calls POST /v1/video-captions with that URL and its own Idempotency-Key, and the agent (or your code) polls GET /v1/jobs/:id/status. The caption read tool video-captions_get is on the list, so the agent can still read the caption resource afterward.

If the source clip is silent, which is common for AI video without audio, the caption job fails with caption_no_speech and next_action: use_overlay_captions. Send cues or segments with text, start and end instead. You can send only one of script_text, words, cues and segments.

Do not depend on the unlisted tools

The docs describe the legacy create tools as callable by name for clients that were already running, not as a supported path for new work. An agent that discovers its tools from tools_list will never see them, and a hard-coded name can disappear. Put the caption burn behind your own REST call and treat the MCP tools as the front half of the chain.

What the agent should do instead

If the agent needs captions today, give it the path the docs list. video_trim is a listed tool, and so is video-captions_get for reading a caption resource. The create tool for captions is legacy and not listed, so an agent that expects a video-captions_create call in its tool list will not find one over the hosted MCP server.

The practical route is to let the agent run the trim over MCP, wait with jobs_wait, read the trimmed URL, and then hand that URL to a step that calls POST /v1/video-captions over the API. That step can be a small server of your own or an Agent Completion whose instructions include the call.

Paid and write tools over MCP need an idempotency_key, so the agent must supply one for the trim. Tell it to reuse the same key on a retry.

  • List tools before planning the chain.
  • Trim over MCP, captions over the API.
  • Always send idempotency_key for paid tools.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume