Gemini 3.8 Flash TTS context caching vs Sume transcripts
Gemini 3.8 Flash TTS prices cached input at $0.125 per million tokens until 2026-12-31. Sume sends the full transcript on each job. When caching matters.

Gemini 3.8 Flash TTS lists a context-caching rate of $0.125 per million tokens through 2026-12-31, rising to $0.25 on 2027-01-01 (read 2026-10-07). The cache applies to the text input side, not to the audio you get back, so it only helps if you resend a large shared prefix. Sume TTS jobs take a transcript of 1 to 20,000 characters each time and have no cache setting.
What the Gemini cache price covers
Gemini bills TTS on two sides: text input tokens and audio output tokens. Context caching discounts the repeated input. At $0.50 per million input tokens against $0.125 cached, a long reused prefix costs a quarter as much to resend. But audio output at $9.00 per million tokens through 2026 is the larger line for narration, and the cache does not touch it.
| Line | Through 2026-12-31 | From 2027-01-01 |
|---|---|---|
| Text input | $0.50 | $1.00 |
| Cached input | $0.125 | $0.25 |
| Audio output | $9.00 | $18.00 |
When a cache would ever pay off
Caching pays off when many requests share a long identical front section, for example a style brief repeated before each line. For ordinary voiceover the input is the script itself, and each line differs. A short ad line has so few input tokens that a cache saves fractions of a cent next to the audio output cost.
So before you design around a cache, estimate your input-to-output ratio. Google's page prices audio output at $9.00 per million tokens against $0.50 for text input, so measure your own ratio on a real script.
What Sume does instead
A Sume TTS 1.0 job is character-metered. Each request carries its own transcript (up to 20,000 characters) or a transcript_source reference to an accepted script, never both. There is no prefix to cache, and no shared-context field in the request schema.
What you can reuse is the result. Send an Idempotency-Key and an exact retry returns the same job instead of making a second paid one. Reusing a key for a different payload fails with 409 idempotency_conflict.
For per-character list rates on the Sonic models, read the live catalog at GET /v1/tts-router/models rather than a number in a blog post. The router docs state the catalog rate is the provider list price times 1.25.
- Gemini: budget the audio output line first, the cache second.
- Sume: budget characters per job, and use idempotency keys for retries.
- Keep the artifact URL from the first Sume result instead of regenerating the same line.
Decision rule
If you run thousands of short lines with a long shared instruction block, price Gemini with the cache and check the 2027 doubling. If your lines are self-contained scripts, caching is a rounding error on either side and the choice rests on voice, language and where the audio needs to land. Sume's docs do not describe Gemini as a routable engine, so the comparison is between separate vendor accounts.
Sources
Related posts
More in Pricing
- Gemini 3.8 Flash TTS: batch and flex $4.50, priority $16.20 per 1M
Gemini 3.8 Flash TTS is $9 per 1M audio tokens Standard, $4.50 batch or flex, $16.20 priority until Dec 31, then double. Cost of 1,000 minutes.
- Gemini Omni 4K is $0.375 a second, double 1080p: worth $3.75 a clip?
Omni Flash 1.1 on Sume costs $0.375 a second at 4K, twice 1080p and three times 720p. A 10-second 4K clip is $3.75 against $1.25 at 720p. When it pays.
- GPT Image 2.5 at 4:5 lists 12-14% below 1:1 at the same quality
On Sume, GPT Image 2.5 at 4:5 (1024x1280) lists 12 to 14 percent below a 1024x1024 square at every quality tier. The five-tier table, with billed prices.
- GPT Image 2.5 at $8 and $30 per million tokens vs Sume's cost
OpenAI prices GPT Image 2.5 per token. Sume meters images per image, returns usage.cost in USD and always reports 0 tokens. Here is how to budget.
Written by Sume