Estimate a Gemini TTS bill from seconds of audio (Python, 25 tok/s)

A 20-line Python estimator for Gemini 3.8 Flash and Flash-Lite TTS audio output cost from seconds, at 25 tokens per second and Google's 2026 and 2027 rates.

4 min readSume
All posts

Multiply seconds by 25 to get audio output tokens, then multiply by the price per million tokens and divide by 1,000,000. Google's pricing page lists standard audio output for Gemini 3.8 Flash TTS at $9.00 per million tokens through 2026 and $18.00 from 2027, and for Flash-Lite TTS at $6.00 and $12.00.

The script below does that arithmetic for any duration. It covers audio output on the standard tier only, so add text input tokens separately.

The estimator

Save it as tts_bill.py and pass a number of seconds. It uses only the standard library.

import sys

TOKENS_PER_SECOND = 25
# USD per million audio output tokens, standard tier: (through 2026, from 2027)
RATES = {"flash": (9.00, 18.00), "flash-lite": (6.00, 12.00)}


def audio_cost(seconds: float, model: str, from_2027: bool = False) -> float:
    rate = RATES[model][1 if from_2027 else 0]
    return seconds * TOKENS_PER_SECOND * rate / 1_000_000


if __name__ == "__main__":
    seconds = float(sys.argv[1]) if len(sys.argv) > 1 else 90.0
    print(f"{seconds:g} s = {seconds * TOKENS_PER_SECOND:,.0f} audio tokens")
    for model in RATES:
        for label, flag in (("2026", False), ("2027", True)):
            print(f"{model:>10} {label}: $" + format(audio_cost(seconds, model, flag), ".5f"))

What it prints for common lengths

These values come from running the formula above, not from a separate quote.

Audio output cost, standard tier, USD, computed from Google's published rates (read 2026-10-07)
SecondsAudio tokensFlash 2026Flash 2027Flash-Lite 2026Flash-Lite 2027
10250$0.00225$0.00450$0.00150$0.00300
902,250$0.02025$0.04050$0.01350$0.02700
60015,000$0.13500$0.27000$0.09000$0.18000
3,60090,000$0.81000$1.62000$0.54000$1.08000

Where the estimate can be wrong

It is a planning number. Three things move the real bill.

  • Seconds are the length of the finished audio, so use a first take, not a word count times a guess.
  • Text input tokens are billed too: $0.50 per million through 2026 and $1.00 from 2027 on the standard tier.
  • Batch and flex are half the standard output price, and priority is 1.8 times it, so change RATES if you run those tiers.

Extending the script

Two small changes make it a real budgeting tool. First, read a list of durations from a file and sum them, so a whole season of episodes gets one total. Second, add the text input side: count your script tokens with Google's token counter and price them at $0.50 per million through 2026. Input is small next to the audio at these rates, but it is not zero, and it doubles in 2027 like the output does.

The same question on Sume

Sume TTS 1.0 does not need a duration estimate to price a job. The charge is the character count times $0.0475 per 1,000, rounded up to the cent, and the OpenAPI reference caps a request at 20,000 characters. If you generate on both engines, run the same script once on each and compare the finished durations before you trust either formula for a big batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume