Gemini 3.8 Flash-Lite TTS for scratch VO vs a Sume TTS first cut

Flash-Lite TTS is Google's cheaper speech tier. If you want a scratch read before locking a script, here is how a first cut on Sume TTS 1.0 compares.

4 min readSume
All posts

Gemini 3.8 Flash-Lite TTS is the low-cost speech tier, listed at about $0.0015 per 10 seconds of audio through 2026 (read 2026-10-07). It suits a scratch read where you are still rewriting the script. Sume TTS 1.0 is the better fit once the voice must be a specific avatar's voice and the audio has to land on a Sume timeline.

What Flash-Lite is for

Google's speech guide lists 100+ languages for Flash-Lite against 130+ for Flash and up to two speakers. Output tokens are $6 per million through 2026, doubling to $12 on 2027-01-01. The pitch is volume: many cheap drafts. A scratch read lets an editor judge pacing before anyone pays for final audio.

Flash-Lite TTS and Sume TTS 1.0 for a first cut - Google pages and Sume docs (read 2026-10-07)
QuestionGemini 3.8 Flash-Lite TTSSume TTS 1.0
VoiceCurated or designed voicesYour avatar's ready voice, or a voice id
Billing unitTokens, about $0.0015 per 10 sCharacters, see the catalog
Script lengthModel context limits1 to 20,000 characters per job
Word and sentence timingsNot in the docs I readtimestamps.words and sentence segments
DeliveryUnary WAV or streaming PCMAsync job, mp3, wav or raw

A first cut on Sume

Even if you draft elsewhere, the Sume first cut has one advantage: it returns word timings. You can compare the draft's rhythm with the scene list without a separate alignment step. Keep the default mp3 for review copies, and switch to wav only when you need sentence slices.

  • Use one idempotency key per script version, so a retry does not buy the same audio twice.
  • Keep speed near 1.0 for the review copy. It runs 0.6 to 1.5.
  • Set language explicitly, since leaving it out means English.

The honest trade

A Gemini scratch track will not sound like your Sume avatar, so judge pacing, not timbre. If you only care about rhythm, the cheaper tier is fine. If a client approves on voice, draft on the voice that ships. Sume does not route to Gemini TTS, so the two live in separate bills.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: script-v3-first-cut" \
  -d '{
    "transcript": "Draft three of the launch voiceover.",
    "avatar_handle": "narrator",
    "language": "en",
    "timestamps": {"words": true},
    "segmentation": {"mode": "sentence"}
  }'

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume