Long video understanding: read the transcript, then pull frames
Google's changelog says Gemini fetches transcripts, frames or audio on demand, with 88% fewer tokens. A transcript-first, frames-second flow on Sume inspect.

The cheapest way to understand a long video is to read the transcript first and look at frames only where the transcript points. That is the pattern Google describes: its Gemini API changelog (read 2026-10-02) says that on 2026-09-01 the model began navigating video timelines, requesting transcripts, frames or audio on demand, with 88% fewer tokens for long-form video. You can run the same shape by hand on Sume's video inspect.
Sume's inspect does not answer questions about a video. It returns probe facts, stills and an optional transcript; your own model or a person reads them.
What do I call, in what order?
Two calls, the second only if needed. First transcribe with sentence segments. Then request frames at the timestamps you chose.
| Pass | Request | Returns |
|---|---|---|
| 1. Transcript | transcribe true, segmentation.mode sentence, frames false | probe, text, words, sentence segments |
| 2. Frames | frames {at: [seconds...]}, up to 24 values | stills as media.sume.com images |
| Skim option | frames {fps: n, seek: fast} | stills snapped to keyframes at or before each instant |
What are the limits?
The source must be 1800 seconds or shorter, and a call returns at most 24 stills. The clip must be a media.sume.com video your workspace already holds; import first, since there is no open-internet fetch. Probe and stills are unbilled; the transcript is the billed half at $0.01 per audio minute, per the docs, to be confirmed in GET /v1/catalog.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: inspect-pass-1" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "frames": false, "transcribe": true, "segmentation": {"mode": "sentence"}}'How do I choose timestamps?
Search the sentence segments for the claim, product name or offer you care about, and use the segment start as an at value. With precise seeking (the default) the still matches the instant; with fast it can be earlier by up to one keyframe interval.
Should I trust the 88% figure for my video?
It is Google's statement about its own model. Measure your own: log tokens per video for a full pass and for a transcript-first pass.
What can the two-pass flow not do?
It cannot tell you what is in a frame. Sume returns the image; reading it takes a model that accepts images or a person. It also cannot hear tone, music or sound effects, since the transcript carries speech only. If audio matters, detach the audio track with Sume's audio detach and listen to the file.
Be honest about coverage too. A transcript-first pass finds what is said; a visual-only moment, such as a price tag or a logo that nobody mentions, needs a frame sample. Add a low-rate fps pass with seek fast to catch those, then look closer where something shows up.
How do I record what I found?
Keep a table of timestamp, quote, still URL and your note. The sentence segments give you the quote and the start time, the still gives you the evidence, and the table lets a second reviewer check your reading without watching the whole video again.
Sources
Related posts
More in Developers
- Reduce video vendor lock-in: swap the model field, keep the pipeline
Keep the video model id in config, read each model's limits from the catalog, and your pipeline survives a vendor change. A script that lists limits.
- Review voice agent call recordings: STT word timings on Sume
After you ship a Gemini Live or other voice agent, transcribe the recordings with Sume STT: word timings, sentence segments, a 10-minute cap per request.
- Voice API deadlines, October 2026 to February 2027
A calendar of voice and transcription API changes from vendor pages: Gemini TTS price rise, OpenAI transcription shutdown, and the xAI voice alias move.
- Wait for an avatar video job with the SDK: waitForJob and its timeout
Avatar jobs need waitForJob, not waitForRun. Defaults, the 20-minute timeout, and why a SumeJobTimeoutError neither cancels nor refunds the render.
Written by Sume