MCP server for video analysis: what can an agent actually read?
Sume's remote MCP lets an agent probe a clip, sample stills, pull exact frames and transcribe audio. Semantic scene tools are dev-only. Limits and costs inside.

An agent can read a clip's facts, stills and speech through Sume's video_inspect and video_frames_create tools. It cannot, on the documented production surface, ask scene-level questions: Sume's docs mark video_analyze and video_segment as development-only, and appearing in tools_list is the condition for calling them. Pick the tool by the question you are asking.
Connect an agent
In Pydantic AI, the MCPToolset class connects a Streamable HTTP server, accepts static headers, and can add a prefix with .prefixed('sume') so tool names do not collide. Its docs call Streamable HTTP the recommended transport for remote servers. Point it at Sume's endpoint and pass your key as a header, or use OAuth through an MCP client that supports it.
Which tool answers which question
| Question | Tool | Documented limit or behaviour |
|---|---|---|
| What is this clip (duration, audio present)? | video_inspect | Reads one clip the workspace already owns; sync mode waits up to 30 seconds, then returns a queued job |
| What does it look like over time? | video_inspect with stills | Source up to 1800 seconds; 24 stills per call; fast can land up to one GOP early, precise is exact |
| Give me this exact frame at full size | video_frames_create, then jobs_wait, then video_frames_get | One of at[] or fps; jpeg default or png; max_edge 16 to 2160 |
| What is said? | video_inspect with transcribe: true | $0.01 per audio minute; a silent clip returns inspect_source_has_no_audio |
Import first
Neither tool fetches from the open internet. The clip must be a media.sume.com artifact or asset in your workspace, so start with media-imports_create. Both inspect and frames are write tools under OAuth and need mcp:write plus an idempotency_key. Frames always returns 202 and needs a poll, while inspect can answer in one call for short clips.
Costs follow compute. The docs say Sume bills probe and stills by their Modal compute, and that each call reserves a compute ceiling at submit, capturing no more than the hold. Transcription adds its per-minute rate to that reservation, and without a duration_seconds hint Sume reserves one minute. For a long clip, pass the hint so the reservation matches the audio, and check the live price in GET /v1/catalog rather than relying on a number copied from a page.
Where semantic questions go
Inspect returns probe facts, stills and optional speech-to-text, not typed scenes. For questions such as which interval shows a product, Sume's docs point to video_analyze and video_segment on the development environment only, gated by configuration. The older video-analyses_* endpoints are being retired, so build new agents on inspect. A practical hybrid is to let the model look at the stills video_inspect returns and reason about them itself, which needs no extra tool.
One refusal to plan for: asking for transcript options such as language_code without transcribe: true returns 400 video_inspect_transcribe_required. Check probe.has_audio first with a stills-free inspect so you never pay to transcribe silence.
Sources
More in Models
- MiniMax H3 first request: a 5-second 768p clip for 38 cents on Sume
A 5-second 768p MiniMax H3 clip costs $0.38 on Sume, with stereo sound included. The request, the 15-second limit and the 2K and 4K upscale prices.
- MiniMax H3 open weights: 768p locally, 2K only hosted
MiniMax's open H3 Base checkpoints top out at 768p. The 2K tier needs a proprietary module. What that means if you want to self-host or call H3 through Sume.
- MiniMax H3 Max 1080p: docs list 768P, Sume lists 1080p
MiniMax's own guide gives H3 Max 480P and 768P. Sume's minimax-h3-max row also lists 1080p. Here is the gap, and what the extra tier means for you.
- MiniMax H3 reference limits: 9 images, 3 videos, 3 audios on Sume
MiniMax H3 reference-to-video on Sume takes up to 9 images, 3 videos and 3 audio clips, 12 in total. The duration rules, the extra-image fee and a request.
Written by Sume