Gemini 3.8 Flash TTS prices rise 50% on Jan 1 2027: recompute
Gemini 3.8 Flash TTS lists $9.00 per million audio tokens until Dec 31, then rises 50%. The new Flash and Flash-Lite numbers, and how to compare with Sume TTS.

Google's pricing page lists Gemini 3.8 Flash TTS at $0.50 per million text input tokens and $9.00 per million audio output tokens through December 31, 2026, with Flash-Lite TTS at $0.50 and $6.00. The same page says pricing rises 50% on January 1, 2027. Applying that rise gives $0.75 and $13.50 for Flash, and $0.75 and $9.00 for Flash-Lite. Those post-rise numbers are my arithmetic, not a figure Google prints.
Because the bill is in tokens, the first job is to learn how many audio tokens your own scripts produce. I did not find a token-per-second figure on the pages I read, so measure it from the usage data of a test call.
The price table
Standard prices are as read on Google's pricing page. The post-rise columns are computed by adding 50%.
- The model page lists Batch, Flex and Priority tiers and caching; Batch is described as 50% off.
- Input limit is 8,192 tokens and output limit 16,384 tokens.
- There is no Live API and no function calling for these TTS models.
| Model | Text input now | Audio output now | Text input from Jan 1 | Audio output from Jan 1 |
|---|---|---|---|---|
| gemini-3.8-flash-tts | $0.50 | $9.00 | $0.75 | $13.50 |
| gemini-3.8-flash-lite-tts | $0.50 | $6.00 | $0.75 | $9.00 |
Recompute a voiceover in four steps
- Run three representative scripts and record the audio output tokens from the usage data.
- Divide the total tokens by the total characters to get your tokens-per-character ratio.
- Multiply expected monthly characters by that ratio, divide by one million, and multiply by the rate.
- Run the same sum with the January rate, and compare with Sume's flat per-character price.
import os, requests
r = requests.post(
"https://api.sume.com/v1/tts-router/generate",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "tts-demo-001",
},
json={
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare three prices.",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"timestamps": {"words": True},
},
timeout=30,
)
r.raise_for_status()
print(r.json())Sume's side of the comparison
Sume bills TTS by character, so the sum needs no token ratio: about $0.0475 per 1,000 characters on Sonic models, rounded up to the cent for each job, as the TTS Router catalog lists. That price is stable across scripts; a token price is not.
What the discount tiers change
Google's model page lists Batch, Flex and Priority tiers. Batch is described as 50% off the standard rate, which helps for work that can wait. If your voiceovers are produced overnight for next-day publishing, ask whether the batch tier fits before you change vendors.
The caching option matters for repeated text, such as a fixed intro line that you speak in every video. Whether it applies to audio output is something to test, since the pages I read describe caching as supported but give no worked example for TTS.
What Sume does not do
Sume does not offer Gemini TTS, inline expression tags such as laugh or sigh, or voice replication from a consent recording. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. For the tags on Sume, you control speed, volume and an optional emotion guide through generation_config instead.
Sources
Related posts
More in Pricing
- Omni Flash 1.1 price chart: 3 to 10 seconds at 4 resolutions
Every whole-second length from 3 to 10 seconds at 360p, 720p, 1080p and 4K for Gemini Omni Flash 1.1 on Sume, from the estimator that reserves your balance.
- Gemini Omni Flash 1.1: what a 10-second 4K clip costs on Sume
A 10-second 4K Gemini Omni Flash 1.1 clip costs $3.75 on Sume; at 1080p it is $1.875 and at 360p $0.375. 3 to 10 seconds, 16:9 or 9:16, audio always on.
- Gemini Omni Flash $17.50 per million tokens vs Sume per-second price
Google lists Gemini Omni Flash video at $17.50 per million tokens, about $0.10 a second at 720p. Thirty 8-second clips: $24.00 at that rate, $30.00 on Sume.
- Gemini Omni Flash 1.1 edit mode: the default reserve is 8 s, $1.00
Editing a clip with Gemini Omni Flash 1.1 on Sume reserves 8 seconds by default: $1.00 at 720p. Send a duration hint of 6 and the reserve drops to $0.75.
Written by Sume