Gemini video understanding API vs Sume video inspect stills

Gemini's API now has agentic video understanding. Sume's video inspect is narrower: probe facts, 8 stills and optional STT you pass to your own model.

4 min readSume
All posts

The Gemini API's video understanding is a model call: Google's changelog entry of September 1, 2026 says the model navigates a video timeline itself. Sume does not offer that. Sume's POST /v1/video-inspect returns probe facts, stills and an optional transcript, and you hand those to whatever model you choose.

Google's side is from its changelog, Sume's from the Video inspect docs, both read 2026-09-30.

What did Google change in the Gemini API?

The September 1 changelog entry, "Agentic video understanding", covers Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite on the Interactions and GenerateContent APIs. The model dynamically navigates video timelines, requesting transcripts, frames or audio tracks on demand. Google states this uses up to 88% fewer tokens for long-form content than static processing. That figure is Google's claim about its own feature; this post did not measure it.

What does Sume video inspect return?

Inspect reads one media.sume.com clip already owned by your workspace and never re-encodes it. The docs say it returns probe facts, stills and optional speech-to-text. It does not produce typed scenes or answer questions about the clip.

Video inspect fields from the Sume docs, read 2026-09-30.
PartWhat the docs say
ProbeReturned as probe; unbilled
StillsOmitted frames gives 8 mid-bin stills; frames: false gives probe only
Chosen stills{ at: [seconds] } with 1 to 24 values, or { fps: n }
TranscriptOnly with transcribe: true: text, words[], optional segments[]
Inputvideo_url on media.sume.com; no open-internet fetch

How do the two approaches differ in practice?

With Gemini, the model decides which frames and audio it needs. With inspect, you decide: pick timestamps with frames.at, or take the default 8 stills, and pass the images and transcript to your own model. Inspect does not answer a question; it supplies evidence. Importing the clip first is required, because the docs state there is no open-internet fetch: use POST /v1/media-imports.

When is inspect the right step?

Use it when you want repeatable inputs for a model you already run, or to check a generated clip's duration and dimensions before a paid step. The docs mark probe and stills as unbilled. For seeking behavior and keyframe stills, see video inspect fast seek.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume