Gemini 3.8 Flash TTS vs Flash-Lite TTS: what one hour of audio costs
One hour of speech is 90,000 audio tokens. Flash TTS is $0.81 and Flash-Lite TTS $0.54 at the 2026 standard rate, doubling in 2027. All tiers compared.

One hour of generated speech is 90,000 audio tokens, because Google says audio tokens correspond to 25 tokens per second and an hour is 3,600 seconds. At the 2026 standard rate that costs $0.81 on Gemini 3.8 Flash TTS ($9.00 per million) and $0.54 on Gemini 3.8 Flash-Lite TTS ($6.00 per million), audio output only.
The gap is one third of the Flash price. Whether that is worth the smaller model depends on languages and on how much acting the script needs, not on the cents.
Cost of one hour of audio, tier by tier
The table multiplies Google's published output price per million audio tokens by 0.09, which is 90,000 tokens as a share of a million. Text input tokens come on top at the listed input price, and I have left them out because they depend on your script length.
| Tier | Flash TTS to Dec 31, 2026 | Flash-Lite TTS to Dec 31, 2026 | Flash TTS from Jan 1, 2027 | Flash-Lite TTS from Jan 1, 2027 |
|---|---|---|---|---|
| Standard | $0.81 | $0.54 | $1.62 | $1.08 |
| Batch / Flex | $0.405 | $0.27 | $0.81 | $0.54 |
| Priority | $1.458 | $0.972 | $2.916 | $1.944 |
What else differs besides price
The speech generation guide gives the model ids gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. It says Flash supports over 130 languages and Flash-Lite over 100. The release notes say Flash-Lite replaces gemini-3.1-flash-tts-preview. The guide also says single-request multi-speaker generation supports up to 2 speakers using prebuilt voices. A scene with three voices therefore needs more than one request.
Picking by job, not by price
A short rule set is enough.
- Use Flash-Lite for bulk narration, scratch tracks and anything a person will hear once at normal speed.
- Use Flash when your language is outside the Flash-Lite list, or when a side-by-side listen shows it acts the script better.
- Use batch or flex if nobody is waiting for the file, since those tiers are half the standard output price in both models.
- Run the same 60 seconds through both before you pick. A minute costs a cent or two, which is less than the time you would spend arguing about it.
Putting a Sume job beside it
Sume TTS 1.0 charges by character, not by token: $0.0475 per 1,000 characters, rounded up to the cent per job. That is not directly comparable with a per-token meter unless you know how many characters your voice reads per second, and Sume does not publish a conversion. Measure it once on your own script: submit the job, read the duration of the finished file, and divide your character count by it.
With that number you can put the Gemini hour and a Sume hour side by side for your own text. Remember the 20,000-character cap per Sume request, and that synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded. The OpenAPI contract states both.
Sources
Related posts
More in Comparisons
- Gemini API leads with Omni Flash, Veo 3.1 for specialists: and Sume?
Google's video docs recommend Gemini Omni Flash as the default and Veo 3.1 for specific needs. Sume has sume/auto or a pinned catalog id. How they line up.
- Gemini Omni Flash 1.1 price: Runway vs Google vs Sume per second
Gemini Omni Flash video costs about $0.10 a second at 720p on Google's page and 10 credits ($0.10) on Runway. Sume sells it at $0.125 before the 5.5% fee.
- Three-second bumper: Gemini Omni Flash 1.1 or Wan 3.0 on Sume
A 3-second bumper costs $0.38 at 720p on Omni Flash 1.1 and $0.38 on Wan 3.0. Omni starts at 3 seconds; Wan starts at 2.
- Grok Imagine, Qwen Image and Imagen 4 Fast all cost 2.5 cents on Sume
Grok Imagine, Qwen Image and Imagen 4 Fast each bill $0.025 per image on Sume. What separates them: ratios, n, edits. Choose the cheap row that fits your job.
Written by Sume