Gemini Omni Flash: 5,792 tokens a second is about $0.10 a second

Google bills Gemini Omni Flash video at $17.50 per million tokens. At 5,792 tokens a second for 720p, that is about $0.10 a second. The math, by clip length.

4 min readSume
All posts

Gemini Omni Flash video output is priced at $17.50 per million tokens, and the Gemini API pricing page lists 5,792 tokens per second of 720p video. Multiply them and one second costs about $0.1014, which the page rounds to about $0.10 per second.

Where the number comes from

The page names gemini-omni-1.1-flash as the GA model on the paid tier, and says the preview id gemini-omni-flash-preview is priced the same. Sume's catalog id is different: gemini-omni-flash-1.1, listed in the video docs with 3 to 10 second clips at 360p, 720p, 1080p, and 4K in 16:9 or 9:16.

Clip cost by length at 720p

Seconds times 5,792 gives tokens, and tokens times $17.50 divided by one million gives dollars. All rows use the 720p token rate only, since this research read did not capture token counts for other resolutions.

Gemini Omni Flash 1.1 at 720p, computed from Google's rates (read 2026-10-03)
Clip lengthTokensProvider cost
3 s17,376$0.304
5 s28,960$0.507
8 s46,336$0.811
10 s57,920$1.014

What it means on Sume

Provider list is not what a Sume job costs. The Sume docs say the workspace balance is reserved at provider list times 1.25 for every model, and usage.cost on the poll response is the billable amount. On that rule, a 10 second 720p clip would reserve about $1.27 before rounding. Treat that as an estimate and read the live pricing_skus and usage.cost for the real figure.

TOKENS_PER_SECOND_720P = 5792
USD_PER_MILLION_TOKENS = 17.50
SUME_MULTIPLIER = 1.25

for seconds in (3, 5, 8, 10):
    tokens = seconds * TOKENS_PER_SECOND_720P
    provider = tokens * USD_PER_MILLION_TOKENS / 1_000_000
    print(f"{seconds:>2} s  {tokens:>6} tokens  provider ${provider:.3f}  estimate with 1.25x ${provider * SUME_MULTIPLIER:.3f}")

Checking the figure yourself

The simplest check is a single paid request. Generate one 5 second clip at 720p, read the usage reported on the response, and compare it with the 28,960 tokens the table implies for 5 seconds. If the numbers disagree, trust the invoice and update your constants.

Token counts for other resolutions are not in the research read, so do not extrapolate the 720p figure to 1080p or 4K. The Gemini pricing page is the place to look for them. On Sume, the same discipline applies to pricing_skus and usage.cost: the first is a rate card, the second is what a job actually billed.

  • Run one pilot clip at the resolution you plan to ship.
  • Compare reported usage with the computed tokens.
  • Store the real per-second cost in your cost model.
  • Re-check when Google changes the model id or tier.

What to take from it

Token pricing hides one useful fact: cost scales with seconds, not with prompt length. A 40 word prompt and a 400 word prompt cost the same for the same clip. What moves the bill is duration, resolution, and how many takes you discard.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume