Long video understanding: read the transcript, then pull frames

Google's changelog says Gemini fetches transcripts, frames or audio on demand, with 88% fewer tokens. A transcript-first, frames-second flow on Sume inspect.

5 min readSume
All posts

The cheapest way to understand a long video is to read the transcript first and look at frames only where the transcript points. That is the pattern Google describes: its Gemini API changelog (read 2026-10-02) says that on 2026-09-01 the model began navigating video timelines, requesting transcripts, frames or audio on demand, with 88% fewer tokens for long-form video. You can run the same shape by hand on Sume's video inspect.

Sume's inspect does not answer questions about a video. It returns probe facts, stills and an optional transcript; your own model or a person reads them.

What do I call, in what order?

Two calls, the second only if needed. First transcribe with sentence segments. Then request frames at the timestamps you chose.

Two-pass inspect. Sume Video inspect docs, read 2026-10-02.
PassRequestReturns
1. Transcripttranscribe true, segmentation.mode sentence, frames falseprobe, text, words, sentence segments
2. Framesframes {at: [seconds...]}, up to 24 valuesstills as media.sume.com images
Skim optionframes {fps: n, seek: fast}stills snapped to keyframes at or before each instant

What are the limits?

The source must be 1800 seconds or shorter, and a call returns at most 24 stills. The clip must be a media.sume.com video your workspace already holds; import first, since there is no open-internet fetch. Probe and stills are unbilled; the transcript is the billed half at $0.01 per audio minute, per the docs, to be confirmed in GET /v1/catalog.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: inspect-pass-1" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "frames": false, "transcribe": true, "segmentation": {"mode": "sentence"}}'

How do I choose timestamps?

Search the sentence segments for the claim, product name or offer you care about, and use the segment start as an at value. With precise seeking (the default) the still matches the instant; with fast it can be earlier by up to one keyframe interval.

Should I trust the 88% figure for my video?

It is Google's statement about its own model. Measure your own: log tokens per video for a full pass and for a transcript-first pass.

What can the two-pass flow not do?

It cannot tell you what is in a frame. Sume returns the image; reading it takes a model that accepts images or a person. It also cannot hear tone, music or sound effects, since the transcript carries speech only. If audio matters, detach the audio track with Sume's audio detach and listen to the file.

Be honest about coverage too. A transcript-first pass finds what is said; a visual-only moment, such as a price tag or a logo that nobody mentions, needs a frame sample. Add a low-rate fps pass with seek fast to catch those, then look closer where something shows up.

How do I record what I found?

Keep a table of timestamp, quote, still URL and your note. The sentence segments give you the quote and the start time, the still gives you the evidence, and the table lets a second reviewer check your reading without watching the whole video again.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume