Gemini Omni Flash: 5,792 tokens a second is about $0.10 a second
Google bills Gemini Omni Flash video at $17.50 per million tokens. At 5,792 tokens a second for 720p, that is about $0.10 a second. The math, by clip length.

Gemini Omni Flash video output is priced at $17.50 per million tokens, and the Gemini API pricing page lists 5,792 tokens per second of 720p video. Multiply them and one second costs about $0.1014, which the page rounds to about $0.10 per second.
Where the number comes from
The page names gemini-omni-1.1-flash as the GA model on the paid tier, and says the preview id gemini-omni-flash-preview is priced the same. Sume's catalog id is different: gemini-omni-flash-1.1, listed in the video docs with 3 to 10 second clips at 360p, 720p, 1080p, and 4K in 16:9 or 9:16.
Clip cost by length at 720p
Seconds times 5,792 gives tokens, and tokens times $17.50 divided by one million gives dollars. All rows use the 720p token rate only, since this research read did not capture token counts for other resolutions.
| Clip length | Tokens | Provider cost |
|---|---|---|
| 3 s | 17,376 | $0.304 |
| 5 s | 28,960 | $0.507 |
| 8 s | 46,336 | $0.811 |
| 10 s | 57,920 | $1.014 |
What it means on Sume
Provider list is not what a Sume job costs. The Sume docs say the workspace balance is reserved at provider list times 1.25 for every model, and usage.cost on the poll response is the billable amount. On that rule, a 10 second 720p clip would reserve about $1.27 before rounding. Treat that as an estimate and read the live pricing_skus and usage.cost for the real figure.
TOKENS_PER_SECOND_720P = 5792
USD_PER_MILLION_TOKENS = 17.50
SUME_MULTIPLIER = 1.25
for seconds in (3, 5, 8, 10):
tokens = seconds * TOKENS_PER_SECOND_720P
provider = tokens * USD_PER_MILLION_TOKENS / 1_000_000
print(f"{seconds:>2} s {tokens:>6} tokens provider ${provider:.3f} estimate with 1.25x ${provider * SUME_MULTIPLIER:.3f}")Checking the figure yourself
The simplest check is a single paid request. Generate one 5 second clip at 720p, read the usage reported on the response, and compare it with the 28,960 tokens the table implies for 5 seconds. If the numbers disagree, trust the invoice and update your constants.
Token counts for other resolutions are not in the research read, so do not extrapolate the 720p figure to 1080p or 4K. The Gemini pricing page is the place to look for them. On Sume, the same discipline applies to pricing_skus and usage.cost: the first is a rate card, the second is what a job actually billed.
- Run one pilot clip at the resolution you plan to ship.
- Compare reported usage with the computed tokens.
- Store the real per-second cost in your cost model.
- Re-check when Google changes the model id or tier.
What to take from it
Token pricing hides one useful fact: cost scales with seconds, not with prompt length. A 40 word prompt and a 400 word prompt cost the same for the same clip. What moves the bill is duration, resolution, and how many takes you discard.
Sources
Related posts
More in Pricing
- GPT Image 2.5 quality auto reserves max on Sume: pin quality
Sume's image docs say quality auto reserves max, and auto size reserves the output token upper bound. What to set instead, plus the 1024 by 1024 figures.
- Grok Imagine video: cost of a 10-second clip on each tier
xAI lists grok-imagine-video-1.5 at $0.080 per second, 1.5-lite at $0.020 and the base model at $0.050. The 10-second math and a draft-then-final plan.
- Is the Veo 3.1 API free? The per-second prices Google lists
Veo 3.1 on the Gemini API is priced per second for Standard, Fast and Lite, charged only on success. The rates, plus 8-second clip totals and what to verify.
- LTX 2.5 Pro at $0.17 a second: 10-second cost and open weights
Higgsfield lists LTX 2.5 Pro at $0.17 per second, so 10 seconds is $1.70. How that compares with running LTX locally and what Sume's catalog lists.
Written by Sume