Gemini Flash-Lite TTS for dubbing: the model is one step of five

Google pitches Gemini 3.8 Flash-Lite TTS for audio dubbing. TTS is one step; detach, transcribe, translate and rejoin are the others. Costs per step.

5 min readSume
All posts

Google says Gemini 3.8 Flash-Lite TTS is optimized for high-volume dubbing, audio content creation and expressive voice agents, but the speech model is only one of five dubbing steps. You still need to pull the audio, transcribe it, translate it and put the new track back.

Where the model sits

A dubbing job has these steps:

  • Detach the audio track from the video.
  • Transcribe it to text with timings.
  • Translate the text, and check length.
  • Synthesize speech in the target language.
  • Join the pieces and replace the audio on the video.

Costs per step

Google lists Gemini 3.8 Flash-Lite TTS audio output at $6.00 per million tokens through December 31, 2026 and $12.00 from January 1, 2027, plus text input at $0.50 and then $1.00 per million tokens. On Sume, the surrounding steps have public rates: audio detach at $0.01 per job, transcription at $0.01 per audio minute, and timeline audio at $0.01 flat per job.

Pieces of a dub (read 2026-10-07)
StepRateSource
Detach audio$0.01 per jobSume rate card
Transcribe$0.01 per audio minuteSume rate card
Synthesize, Flash-Lite TTS$6.00 per 1M audio tokens until Dec 31, 2026Google pricing
Join with timeline audio$0.01 flat per jobSume rate card

Timing is the hard part

A translated line is often longer or shorter than the original. Use sentence segments from the transcript to keep each line inside its original window, and rewrite lines that run long instead of speeding them up. Sume's STT returns sentence segments with start and end times when you ask for sentence segmentation.

Keep WAV between steps. Sume's timeline audio doc notes that MP3 adds priming padding at each edge, so use WAV when you will join the file again.

Pick the speech model last

Decide the pipeline and the timing rules first, then audition speech models on three real lines. The model that wins on one language may not win on another.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume