Estimate a Gemini TTS bill from seconds of audio (Python, 25 tok/s)
A 20-line Python estimator for Gemini 3.8 Flash and Flash-Lite TTS audio output cost from seconds, at 25 tokens per second and Google's 2026 and 2027 rates.

Multiply seconds by 25 to get audio output tokens, then multiply by the price per million tokens and divide by 1,000,000. Google's pricing page lists standard audio output for Gemini 3.8 Flash TTS at $9.00 per million tokens through 2026 and $18.00 from 2027, and for Flash-Lite TTS at $6.00 and $12.00.
The script below does that arithmetic for any duration. It covers audio output on the standard tier only, so add text input tokens separately.
The estimator
Save it as tts_bill.py and pass a number of seconds. It uses only the standard library.
import sys
TOKENS_PER_SECOND = 25
# USD per million audio output tokens, standard tier: (through 2026, from 2027)
RATES = {"flash": (9.00, 18.00), "flash-lite": (6.00, 12.00)}
def audio_cost(seconds: float, model: str, from_2027: bool = False) -> float:
rate = RATES[model][1 if from_2027 else 0]
return seconds * TOKENS_PER_SECOND * rate / 1_000_000
if __name__ == "__main__":
seconds = float(sys.argv[1]) if len(sys.argv) > 1 else 90.0
print(f"{seconds:g} s = {seconds * TOKENS_PER_SECOND:,.0f} audio tokens")
for model in RATES:
for label, flag in (("2026", False), ("2027", True)):
print(f"{model:>10} {label}: $" + format(audio_cost(seconds, model, flag), ".5f"))What it prints for common lengths
These values come from running the formula above, not from a separate quote.
| Seconds | Audio tokens | Flash 2026 | Flash 2027 | Flash-Lite 2026 | Flash-Lite 2027 |
|---|---|---|---|---|---|
| 10 | 250 | $0.00225 | $0.00450 | $0.00150 | $0.00300 |
| 90 | 2,250 | $0.02025 | $0.04050 | $0.01350 | $0.02700 |
| 600 | 15,000 | $0.13500 | $0.27000 | $0.09000 | $0.18000 |
| 3,600 | 90,000 | $0.81000 | $1.62000 | $0.54000 | $1.08000 |
Where the estimate can be wrong
It is a planning number. Three things move the real bill.
- Seconds are the length of the finished audio, so use a first take, not a word count times a guess.
- Text input tokens are billed too: $0.50 per million through 2026 and $1.00 from 2027 on the standard tier.
- Batch and flex are half the standard output price, and priority is 1.8 times it, so change
RATESif you run those tiers.
Extending the script
Two small changes make it a real budgeting tool. First, read a list of durations from a file and sum them, so a whole season of episodes gets one total. Second, add the text input side: count your script tokens with Google's token counter and price them at $0.50 per million through 2026. Input is small next to the audio at these rates, but it is not zero, and it doubles in 2027 like the output does.
The same question on Sume
Sume TTS 1.0 does not need a duration estimate to price a job. The charge is the character count times $0.0475 per 1,000, rounded up to the cent, and the OpenAPI reference caps a request at 20,000 characters. If you generate on both engines, run the same script once on each and compare the finished durations before you trust either formula for a big batch.
Sources
Related posts
More in Developers
- Fade in and out on a Timeline render: output fade seconds 0 to 5
Set output.fade_in_seconds and fade_out_seconds (0 to 5 s, sum within the render length). The music bed has its own fade_out_seconds, up to 10.
- Fast-cut Shorts in Timeline: eight chained fades, then a hard cut
Timeline refuses more than 8 adjacent fades with too_many_chained_transitions. Transitions must be 1 s or less and half the shorter neighbour. How to plan cuts.
- FastMCP 4 client OAuth with Sume: request mcp:read, add write later
Use FastMCP's OAuth helper in Python to sign in to Sume's hosted MCP with mcp:read first, then ask for mcp:write only for the script that spends.
- Fix an underexposed photo: curves first, AI edit only if needed
Underexposed photo? Try Pillow autocontrast and gamma for free, then an ideogram/ideogram-v4.5 edit at $0.075 only if noise or colour needs more.
Written by Sume