Google Vids AI voiceover: the same workflow through an API

Google Vids is reported to have more natural voiceovers on Gemini 3.8 Flash Lite TTS. To script a voiceover outside Vids, use a Sume TTS job plus Timeline.

4 min readSume
All posts

Google Vids is a point-and-click editor. If you want the same voiceover step in code, write the script, generate speech with a Sume TTS job, and lay the audio under your clips in Timeline 1.0. That gives you a repeatable pipeline with job ids, not a session in an editor.

Releasebot's Google page (read 2026-10-02) lists the 2026-10-02 Workspace weekly recap, which mentions more natural voiceovers from an upgraded Gemini 3.8 Flash Lite TTS in Google Vids, and customizable captions. Releasebot is a third-party aggregator, so read this as reported, not as Google's own wording. Sume's TTS is a different product with its own voices; see the engine note below.

What are the steps?

Four calls cover it. The table lists each, with the Sume surface.

Voiceover steps outside Vids. Sume docs, read 2026-10-02; Vids facts from Releasebot, read 2026-10-02.
StepSume surfaceNote
Write the scriptYour own textUp to 20,000 characters per TTS request
Speak itPOST /v1/tts-1.0/generateVoice from avatar_id, avatar_handle or voice.id
Join takesPOST /v1/timeline-1.0/audio, operation concat1 to 20 parts, no re-synthesis
Place under videoPOST /v1/timeline-1.0/renderaudio.url or audio.parts[] as the spine

Which engine does Sume's TTS use?

Sume's TTS 1.0 has no engine picker: model and model_id are rejected with 400. The separate TTS Router lists explicit models at GET /v1/tts-router/models. Check that catalog for what is routable today rather than assuming it matches an engine named in Google's recap.

What output format should I ask for?

The default is mp3 at 44100 Hz and 128 kbps. When the audio will feed an avatar mux, the OpenAPI notes say to pass wav with pcm_s16le at 44100 explicitly. For a plain video voiceover the default is fine.

Can Sume match Vids captions?

Sume has its own caption styles through POST /v1/video-captions; see the companion post on aligning captions to a script. It does not copy Vids' caption options.

Where does a scripted pipeline beat an editor?

An editor is quick for one video. A pipeline pays off at the tenth: the script is a file, the voice is a parameter, and the render is a job with an id. When a price or a claim changes, you edit one line and re-run, instead of reopening a project and re-recording.

It also helps review. A script in version control has a history, and each voiceover job records the transcript it was generated from. A reviewer reads the text, and the audio follows from it.

What does a minimal run look like?

Generate each paragraph as its own TTS job, keep the ids in order, join the audio with a concat job, and send the joined file to Timeline as the audio spine. If a paragraph reads badly, regenerate only that one and join again. The sentence segmentation option can return gapless per-sentence slices when you need to line visuals up with sentences.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume