Gemini 3.8 Flash-Lite TTS for scratch VO vs a Sume TTS first cut
Flash-Lite TTS is Google's cheaper speech tier. If you want a scratch read before locking a script, here is how a first cut on Sume TTS 1.0 compares.

Gemini 3.8 Flash-Lite TTS is the low-cost speech tier, listed at about $0.0015 per 10 seconds of audio through 2026 (read 2026-10-07). It suits a scratch read where you are still rewriting the script. Sume TTS 1.0 is the better fit once the voice must be a specific avatar's voice and the audio has to land on a Sume timeline.
What Flash-Lite is for
Google's speech guide lists 100+ languages for Flash-Lite against 130+ for Flash and up to two speakers. Output tokens are $6 per million through 2026, doubling to $12 on 2027-01-01. The pitch is volume: many cheap drafts. A scratch read lets an editor judge pacing before anyone pays for final audio.
| Question | Gemini 3.8 Flash-Lite TTS | Sume TTS 1.0 |
|---|---|---|
| Voice | Curated or designed voices | Your avatar's ready voice, or a voice id |
| Billing unit | Tokens, about $0.0015 per 10 s | Characters, see the catalog |
| Script length | Model context limits | 1 to 20,000 characters per job |
| Word and sentence timings | Not in the docs I read | timestamps.words and sentence segments |
| Delivery | Unary WAV or streaming PCM | Async job, mp3, wav or raw |
A first cut on Sume
Even if you draft elsewhere, the Sume first cut has one advantage: it returns word timings. You can compare the draft's rhythm with the scene list without a separate alignment step. Keep the default mp3 for review copies, and switch to wav only when you need sentence slices.
- Use one idempotency key per script version, so a retry does not buy the same audio twice.
- Keep
speednear 1.0 for the review copy. It runs 0.6 to 1.5. - Set
languageexplicitly, since leaving it out means English.
The honest trade
A Gemini scratch track will not sound like your Sume avatar, so judge pacing, not timbre. If you only care about rhythm, the cheaper tier is fine. If a client approves on voice, draft on the voice that ships. Sume does not route to Gemini TTS, so the two live in separate bills.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: script-v3-first-cut" \
-d '{
"transcript": "Draft three of the launch voiceover.",
"avatar_handle": "narrator",
"language": "en",
"timestamps": {"words": true},
"segmentation": {"mode": "sentence"}
}'Sources
Related posts
More in Comparisons
- Gemini 3.8 Flash TTS voice library vs Sume voice selectors
Gemini 3.8 Flash TTS lists 30 curated voices plus a larger extended library. Sume picks a voice by avatar or voice id. How the two selection models differ.
- Gemini API leads with Omni Flash, Veo 3.1 for specialists: and Sume?
Google's video docs recommend Gemini Omni Flash as the default and Veo 3.1 for specific needs. Sume has sume/auto or a pinned catalog id. How they line up.
- Grok Imagine, Qwen Image and Imagen 4 Fast all cost 2.5 cents on Sume
Grok Imagine, Qwen Image and Imagen 4 Fast each bill $0.025 per image on Sume. What separates them: ratios, n, edits. Choose the cheap row that fits your job.
- HappyHorse 1.0 vs 1.1: which one renders 480P on Model Studio?
On Alibaba Model Studio only HappyHorse 1.1 lists 480P. Both versions take 3 to 15 seconds, and 1.0 adds a video-edit id. Neither is in Sume's catalog.
Written by Sume