Gemini 3.8 Flash TTS for video voiceover: what Sume TTS returns

Gemini 3.8 Flash TTS takes text and returns speech. For a video voiceover you also need word times and clip cuts. Here is what a Sume TTS job adds.

5 min readSume
All posts

Gemini 3.8 Flash TTS turns text into speech, and its speech generation guide says these models accept text-only input. A video voiceover needs more than audio: word times for captions and sentence cuts for scenes. A Sume TTS 1.0 job returns those in the same result, for $0.0475 per 1,000 characters.

What the Gemini page says

The guide lists gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. Both support voice design and voice replication, Flash covers over 130 languages and Flash-Lite over 100, and multi-speaker output supports up to two speakers with prebuilt voices. Input is text only.

That page describes the speech you get back. It is a speech model, and a video edit needs a second step to find where each word falls.

What a Sume TTS job returns

The result carries audio_url, model_id, the resolved voice, language, output_format and the echoed generation_config. Two optional blocks add the editing data:

  • timestamps: {words: true} returns words[] with start and end times.
  • segmentation: {mode: "sentence"} returns segments[]; wav or raw output also gives each sentence its own audio slice. It requires word timestamps, otherwise the API answers 400 segmentation_requires_word_timestamps.
  • boundary_lead_ms (0 to 500, default 70) moves each cut earlier so a consonant is not clipped.

The editing recipe

Request wav, pcm_s16le, 44100 Hz, words on, sentence segmentation on. You get one file for the whole take and a slice per sentence. Use the slice lengths to set scene durations, and send the words to video captions so the burned-in text lines up with the speech.

The transcript receipt on the result proves which text was submitted. It does not prove how a name was pronounced, so listen to the take before it goes into a render.

Choosing between them

Pick the Gemini route when you already live in the Gemini API and only need audio. Pick Sume TTS when the audio feeds a video: Sume's voices are Cartesia Sonic 3.6, the output is a durable media URL, and the same job gives you the timing data. If you liked the voice replication idea from Gemini 3.8, a Sume voice clone gives you a voice.id you can reuse in every job.

A quick pre-flight

Before you commit a script, run a dry_run over MCP or read the price from the character count: characters divided by 1,000, times $0.0475. A 20,000-character request is the maximum and costs $0.95. If the job fails the language check, fix the language code rather than forcing the voice.

Then render once, read words[], and confirm the last word's end time fits the video slot you have.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume