Gemini video understanding: 100 vs 300 tokens a second vs Sume

Gemini reads video at about 100 tokens a second, or 300 at high resolution: 60,000 or 180,000 tokens for 10 minutes. How Sume's stills route compares.

4 min readSume
All posts

Gemini's video understanding docs say a video costs about 100 tokens per second at the default (low) media resolution and about 300 tokens per second at high media resolution, sampled at 1 frame per second. A 10-minute clip is therefore about 60,000 tokens at default and 180,000 at high. Sume's video_inspect does something different: it probes a clip and returns up to 24 still images, so you pass those to whichever model you use for reasoning.

Read on 2026-10-01: Gemini API video understanding and Sume's Video inspect docs.

What do the token numbers mean for a clip?

The arithmetic is seconds times rate. Google adds that a model with a 1M-token context window can take up to 3 hours of video at low resolution or 1 hour at high. Pick the rate by how much detail you need: reading on-screen text or small objects is a reason for 300; summarizing a talk is not.

Computed from the rates on Google's page, read 2026-10-01.
Clip lengthDefault (~100/s)High (~300/s)
30 seconds3,0009,000
5 minutes30,00090,000
10 minutes60,000180,000
60 minutes360,0001,080,000

What does Sume's video inspect return?

POST /v1/video-inspect takes one media.sume.com clip already owned by your workspace and returns probe facts (the probe object), stills, and an optional transcript. Probe and stills are unbilled; the transcript is billed at a public rate of $0.01 per audio minute and the docs tell you to confirm it in GET /v1/catalog. With frames omitted you get 8 stills; you can ask for explicit timestamps or an fps, up to 24 stills per call, on a source of up to 1800 seconds. Idempotency-Key is required, and off-host URLs are rejected, so import the file first.

From the Sume docs, read 2026-10-01.
ItemValue
Source cap1800 seconds
Stills per callUp to 24; 8 by default
Frame optionsat[] (1 to 24 timestamps) or fps (0 < n <= 2), not both
Still sizemax_edge 64 to 2160, default 768
TranscriptOptional; $0.01 per audio minute

Which should you use?

They answer different questions. Gemini ingests the whole video as tokens and reasons over motion and sound. Sume's route gives you evidence frames and facts you can attach to a prompt or a review. For a 30-minute source, 24 stills is a sample, not a watch-through. Sume's docs list semantic scene questions as a separate tool, available only when it appears in tools_list, so do not assume it.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume