Gemini video understanding API vs Sume video inspect stills
Gemini's API now has agentic video understanding. Sume's video inspect is narrower: probe facts, 8 stills and optional STT you pass to your own model.

The Gemini API's video understanding is a model call: Google's changelog entry of September 1, 2026 says the model navigates a video timeline itself. Sume does not offer that. Sume's POST /v1/video-inspect returns probe facts, stills and an optional transcript, and you hand those to whatever model you choose.
Google's side is from its changelog, Sume's from the Video inspect docs, both read 2026-09-30.
What did Google change in the Gemini API?
The September 1 changelog entry, "Agentic video understanding", covers Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite on the Interactions and GenerateContent APIs. The model dynamically navigates video timelines, requesting transcripts, frames or audio tracks on demand. Google states this uses up to 88% fewer tokens for long-form content than static processing. That figure is Google's claim about its own feature; this post did not measure it.
What does Sume video inspect return?
Inspect reads one media.sume.com clip already owned by your workspace and never re-encodes it. The docs say it returns probe facts, stills and optional speech-to-text. It does not produce typed scenes or answer questions about the clip.
| Part | What the docs say |
|---|---|
| Probe | Returned as probe; unbilled |
| Stills | Omitted frames gives 8 mid-bin stills; frames: false gives probe only |
| Chosen stills | { at: [seconds] } with 1 to 24 values, or { fps: n } |
| Transcript | Only with transcribe: true: text, words[], optional segments[] |
| Input | video_url on media.sume.com; no open-internet fetch |
How do the two approaches differ in practice?
With Gemini, the model decides which frames and audio it needs. With inspect, you decide: pick timestamps with frames.at, or take the default 8 stills, and pass the images and transcript to your own model. Inspect does not answer a question; it supplies evidence. Importing the clip first is required, because the docs state there is no open-internet fetch: use POST /v1/media-imports.
When is inspect the right step?
Use it when you want repeatable inputs for a model you already run, or to check a generated clip's duration and dimensions before a paid step. The docs mark probe and stills as unbilled. For seeking behavior and keyframe stills, see video inspect fast seek.
Sources
Related posts
More in Developers
- Gemini CLI MCP timeout default vs Sume's 55-second jobs_wait
Gemini CLI's MCP timeout defaults to 600,000 ms. Sume's jobs_wait holds at most 55 seconds per call: repeat it on wait_slice_expired, never resubmit.
- Gemini CLI trust: true on a Sume MCP server: what it skips
Gemini CLI's trust setting bypasses every tool confirmation dialog. What that means for Sume's paid tools, and which Sume gates still apply.
- Gemini Omni first and last frame: image_url + end_image_url
Google calls it Interpolation (first + last frame). On Sume you send image_url plus end_image_url to gemini-omni-flash-1.1 through the Video Router.
- Gemini rate limits: per project, not per API key. And Sume?
Google says Gemini rate limits apply per project, not per API key, so a second key adds nothing. What Sume's docs say about plan limits and how to read them.
Written by Sume