Gemini video understanding: 100 vs 300 tokens a second vs Sume
Gemini reads video at about 100 tokens a second, or 300 at high resolution: 60,000 or 180,000 tokens for 10 minutes. How Sume's stills route compares.

Gemini's video understanding docs say a video costs about 100 tokens per second at the default (low) media resolution and about 300 tokens per second at high media resolution, sampled at 1 frame per second. A 10-minute clip is therefore about 60,000 tokens at default and 180,000 at high. Sume's video_inspect does something different: it probes a clip and returns up to 24 still images, so you pass those to whichever model you use for reasoning.
Read on 2026-10-01: Gemini API video understanding and Sume's Video inspect docs.
What do the token numbers mean for a clip?
The arithmetic is seconds times rate. Google adds that a model with a 1M-token context window can take up to 3 hours of video at low resolution or 1 hour at high. Pick the rate by how much detail you need: reading on-screen text or small objects is a reason for 300; summarizing a talk is not.
| Clip length | Default (~100/s) | High (~300/s) |
|---|---|---|
| 30 seconds | 3,000 | 9,000 |
| 5 minutes | 30,000 | 90,000 |
| 10 minutes | 60,000 | 180,000 |
| 60 minutes | 360,000 | 1,080,000 |
What does Sume's video inspect return?
POST /v1/video-inspect takes one media.sume.com clip already owned by your workspace and returns probe facts (the probe object), stills, and an optional transcript. Probe and stills are unbilled; the transcript is billed at a public rate of $0.01 per audio minute and the docs tell you to confirm it in GET /v1/catalog. With frames omitted you get 8 stills; you can ask for explicit timestamps or an fps, up to 24 stills per call, on a source of up to 1800 seconds. Idempotency-Key is required, and off-host URLs are rejected, so import the file first.
| Item | Value |
|---|---|
| Source cap | 1800 seconds |
| Stills per call | Up to 24; 8 by default |
| Frame options | at[] (1 to 24 timestamps) or fps (0 < n <= 2), not both |
| Still size | max_edge 64 to 2160, default 768 |
| Transcript | Optional; $0.01 per audio minute |
Which should you use?
They answer different questions. Gemini ingests the whole video as tokens and reasons over motion and sound. Sume's route gives you evidence frames and facts you can attach to a prompt or a review. For a 30-minute source, 24 stills is a sample, not a watch-through. Sume's docs list semantic scene questions as a separate tool, available only when it appears in tools_list, so do not assume it.
Sources
Related posts
More in Media tools
- Hook, demo, CTA: assemble a product ad from 3 clips in one render
Build a 15-second holiday product ad from a hook clip, a demo clip and a call-to-action card with one Sume timeline render. Plan first, render second.
- hypit boundaries: cut candidates and the 0.1 threshold
hypit boundaries scores possible cuts from a 12 fps 32x32 difference pass. What rate and threshold do, and how it differs from reference ingest's shot list.
- Word-level transcript with confidence scores: hypit transcribe
hypit transcribe returns each word with start, end and a 0-1 score using WhisperX, or Sume STT 1.0 without scores. Engines, checkpoints, billing and errors.
- Keyframe trim starts early: fix it with actual_start_seconds
A video-trim with precision keyframe can begin a GOP before your start time. The result reports actual_start_seconds; use it to re-base the next step.
Written by Sume