Omni's 1M-token context: how many ten-second edit turns fit?

Omni lists a 1,048,576-token context and 5,792 tokens per video second. If prior clips stay in context, 18 ten-second turns fit; Sume edits are stateless.

4 min readSume
All posts

Google's Omni model page lists a context window of 1,048,576 tokens, and the pricing page counts video output at 5,792 tokens per second of 720p video. If every earlier clip in a multi-turn previous_interaction_id chain stayed in context, that is 57,920 tokens per ten-second clip, so about 18 turns before the window fills. Google's pages do not say whether prior outputs count, so treat 18 as a worst case.

The arithmetic

The inputs are the model page window, the pricing page token rate, and the Omni guide statement that multi-turn editing uses previous_interaction_id. Divide the window by tokens per clip and round down.

Worst-case turns in a 1,048,576-token window at 5,792 tokens per second (read 2026-10-04)
Clip lengthTokens per clipClips that fit
3 seconds17,37660
5 seconds28,96036
8 seconds46,33622
10 seconds57,92018

What the numbers do not tell you

Prompts, reference images and your own text also take tokens, and the 720p rate is the only one I read; other resolutions may count differently. Whether Google keeps old video output in the context, summarizes it, or drops it is not stated, so do not design around the figure; run a long session once and see where it fails.

Also remember the shape of the work: most real edits are three to five turns. A session that needs 18 passes is usually a sign to restart from the best frame.

The stateless alternative

Sume's edit mode has no session. A Video Router edit request carries video_url plus a prompt, cannot be combined with image or reference inputs, and bills output seconds at provider list times 1.25. Each edit is a separate job you poll through Jobs and results.

That means no context ceiling, but also no memory: you restate what to keep in every prompt, and each edit starts from the exact clip you pass. For long chains, keep your own prompt log and pass the last accepted output as the next video_url.

Sources

Related posts

More in Models

All Models posts

Written by Sume