Gemini video understanding 88% fewer tokens vs Sume Video inspect
Gemini reports up to 88% fewer tokens on long video. Sume Video inspect and Reference ingest take another route: stills, transcript and a manifest.

The Gemini API changelog for September 1, 2026 says its agentic video understanding approach uses up to 88% fewer tokens for long-form content than static processing. Sume's two read paths do not use model tokens in that way: Video inspect returns a probe, sampled stills and an optional transcript, and Reference ingest returns a timestamped manifest for one clip.
What Google reports
The changelog entry is the only Gemini claim used here: up to 88% fewer tokens on long-form content compared with static processing. This post does not repeat a benchmark or a price, and the figure is an upper bound, not a typical result.
What Sume returns
Both Sume surfaces read one media.sume.com clip that the workspace already owns.
| Surface | Source cap | Returns |
|---|---|---|
Video inspect (POST /v1/video-inspect) | 1800 s | Probe, stills (8 by default, up to 24 per call), optional transcript |
Reference ingest (POST /v1/reference-ingest) | 300 s | Shots, on-screen text lines, audio facts, scene-change candidates, one labeled overview strip |
How you control the cost
With Video inspect you choose the program. Omit frames for 8 mid-bin stills, pass frames: false for a probe only, pass { at: [...] } for named times or { fps: n } for a rate no higher than 2. seek: "fast" snaps to keyframes when you are only skimming. A transcript is optional and billed: transcribe: true runs Sume STT 1.0 on the audio.
Probe and stills are billed by their Modal compute rather than by tokens. Reference ingest transcribes only when the track is not silent and speech is detected, and settles that line to zero otherwise.
Which to use
If you want a model to answer an open question about a long video, an agentic video-understanding API fits that job. If you want facts and stills to feed an agent or a remix plan, Sume's surfaces give you structured output you can inspect yourself. They are inputs to a model, not a replacement for one. Compare cost on your own clips, because the two are billed on different units.
Sources
Related posts
More in Developers
- Gemini CLI 0.62 MCP titles: reading Sume's tool names
Gemini CLI v0.62.0 formats MCP tool call titles as structured signatures. Sume tool ids are underscore names such as generate_image; dotted aliases map to them.
- Gemini CLI v0.63 plan execution in CI: gate paid Sume calls first
Gemini CLI preview v0.63.0 adds autonomous plan execution in non-interactive mode. Before unattended runs, gate Sume paid tools with dry_run and max_spend_usd.
- Headless Gemini CLI and MCP auth: use a Sume API key
A headless Gemini CLI run cannot finish a browser consent. Connect to hosted Sume MCP with an API key header instead, and keep paid calls bounded.
- Gemini CLI untrusted tool output provenance: Sume results as data
Gemini CLI 0.60 and 0.61 release notes name fixes for untrusted tool output and indirect prompt injection. Keep Sume write tools behind scope and caps.
Written by Sume