Gemini 3.8 Flash TTS for video voiceover: what Sume TTS returns
Gemini 3.8 Flash TTS takes text and returns speech. For a video voiceover you also need word times and clip cuts. Here is what a Sume TTS job adds.

Gemini 3.8 Flash TTS turns text into speech, and its speech generation guide says these models accept text-only input. A video voiceover needs more than audio: word times for captions and sentence cuts for scenes. A Sume TTS 1.0 job returns those in the same result, for $0.0475 per 1,000 characters.
What the Gemini page says
The guide lists gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. Both support voice design and voice replication, Flash covers over 130 languages and Flash-Lite over 100, and multi-speaker output supports up to two speakers with prebuilt voices. Input is text only.
That page describes the speech you get back. It is a speech model, and a video edit needs a second step to find where each word falls.
What a Sume TTS job returns
The result carries audio_url, model_id, the resolved voice, language, output_format and the echoed generation_config. Two optional blocks add the editing data:
timestamps: {words: true}returnswords[]with start and end times.segmentation: {mode: "sentence"}returnssegments[]; wav or raw output also gives each sentence its own audio slice. It requires word timestamps, otherwise the API answers 400segmentation_requires_word_timestamps.boundary_lead_ms(0 to 500, default 70) moves each cut earlier so a consonant is not clipped.
The editing recipe
Request wav, pcm_s16le, 44100 Hz, words on, sentence segmentation on. You get one file for the whole take and a slice per sentence. Use the slice lengths to set scene durations, and send the words to video captions so the burned-in text lines up with the speech.
The transcript receipt on the result proves which text was submitted. It does not prove how a name was pronounced, so listen to the take before it goes into a render.
Choosing between them
Pick the Gemini route when you already live in the Gemini API and only need audio. Pick Sume TTS when the audio feeds a video: Sume's voices are Cartesia Sonic 3.6, the output is a durable media URL, and the same job gives you the timing data. If you liked the voice replication idea from Gemini 3.8, a Sume voice clone gives you a voice.id you can reuse in every job.
A quick pre-flight
Before you commit a script, run a dry_run over MCP or read the price from the character count: characters divided by 1,000, times $0.0475. A 20,000-character request is the maximum and costs $0.95. If the job fails the language check, fix the language code rather than forcing the voice.
Then render once, read words[], and confirm the last word's end time fits the video slot you have.
Sources
Related posts
More in Comparisons
- Voice agent cost per minute: Gemini 3.8 Live vs gpt-realtime-2.1
Gemini 3.8 Live lists $0.005 per minute in and $0.018 out; gpt-realtime-2.1 lists $32 / $64 per 1M audio tokens. Convert tokens to minutes before you compare.
- Google's Omni-first video rule has an expiry: Veo 3.1 ends Oct 22
Google's Gemini API docs name Omni Flash the default for video and keep Veo 3.1 for extension, last-frame control and legacy pipelines until October 22.
- Omni 1.1 Flash 1080p vs Wan 3.0 1080p: price for 30 seconds
A 30 second 1080p film costs $7.50 on Sume's Wan 3.0 in one request, or $5.625 as three 10 s Gemini Omni 1.1 Flash clips. What the cheaper route costs you.
- Omni 1.1 Flash extends to 40 s; Seedance 2.5 makes 30 s in one pass
Google extends Omni 1.1 Flash clips in 10 s steps to 40 s. Seedance 2.5 generates 30 s in one pass. How to get 40 s on Sume with Seedance and Timeline.
Written by Sume