MCP server for video analysis: what can an agent actually read?

Sume's remote MCP lets an agent probe a clip, sample stills, pull exact frames and transcribe audio. Semantic scene tools are dev-only. Limits and costs inside.

5 min readSume
All posts

An agent can read a clip's facts, stills and speech through Sume's video_inspect and video_frames_create tools. It cannot, on the documented production surface, ask scene-level questions: Sume's docs mark video_analyze and video_segment as development-only, and appearing in tools_list is the condition for calling them. Pick the tool by the question you are asking.

Connect an agent

In Pydantic AI, the MCPToolset class connects a Streamable HTTP server, accepts static headers, and can add a prefix with .prefixed('sume') so tool names do not collide. Its docs call Streamable HTTP the recommended transport for remote servers. Point it at Sume's endpoint and pass your key as a header, or use OAuth through an MCP client that supports it.

Which tool answers which question

Sume video reading tools and documented limits (read 2026-10-06)
QuestionToolDocumented limit or behaviour
What is this clip (duration, audio present)?video_inspectReads one clip the workspace already owns; sync mode waits up to 30 seconds, then returns a queued job
What does it look like over time?video_inspect with stillsSource up to 1800 seconds; 24 stills per call; fast can land up to one GOP early, precise is exact
Give me this exact frame at full sizevideo_frames_create, then jobs_wait, then video_frames_getOne of at[] or fps; jpeg default or png; max_edge 16 to 2160
What is said?video_inspect with transcribe: true$0.01 per audio minute; a silent clip returns inspect_source_has_no_audio

Import first

Neither tool fetches from the open internet. The clip must be a media.sume.com artifact or asset in your workspace, so start with media-imports_create. Both inspect and frames are write tools under OAuth and need mcp:write plus an idempotency_key. Frames always returns 202 and needs a poll, while inspect can answer in one call for short clips.

Costs follow compute. The docs say Sume bills probe and stills by their Modal compute, and that each call reserves a compute ceiling at submit, capturing no more than the hold. Transcription adds its per-minute rate to that reservation, and without a duration_seconds hint Sume reserves one minute. For a long clip, pass the hint so the reservation matches the audio, and check the live price in GET /v1/catalog rather than relying on a number copied from a page.

Where semantic questions go

Inspect returns probe facts, stills and optional speech-to-text, not typed scenes. For questions such as which interval shows a product, Sume's docs point to video_analyze and video_segment on the development environment only, gated by configuration. The older video-analyses_* endpoints are being retired, so build new agents on inspect. A practical hybrid is to let the model look at the stills video_inspect returns and reason about them itself, which needs no extra tool.

One refusal to plan for: asking for transcript options such as language_code without transcribe: true returns 400 video_inspect_transcribe_required. Check probe.has_audio first with a stills-free inspect so you never pay to transcribe silence.

Sources

More in Models

All Models posts

Written by Sume