Google Vids AI voiceover: the same workflow through an API
Google Vids is reported to have more natural voiceovers on Gemini 3.8 Flash Lite TTS. To script a voiceover outside Vids, use a Sume TTS job plus Timeline.

Google Vids is a point-and-click editor. If you want the same voiceover step in code, write the script, generate speech with a Sume TTS job, and lay the audio under your clips in Timeline 1.0. That gives you a repeatable pipeline with job ids, not a session in an editor.
Releasebot's Google page (read 2026-10-02) lists the 2026-10-02 Workspace weekly recap, which mentions more natural voiceovers from an upgraded Gemini 3.8 Flash Lite TTS in Google Vids, and customizable captions. Releasebot is a third-party aggregator, so read this as reported, not as Google's own wording. Sume's TTS is a different product with its own voices; see the engine note below.
What are the steps?
Four calls cover it. The table lists each, with the Sume surface.
| Step | Sume surface | Note |
|---|---|---|
| Write the script | Your own text | Up to 20,000 characters per TTS request |
| Speak it | POST /v1/tts-1.0/generate | Voice from avatar_id, avatar_handle or voice.id |
| Join takes | POST /v1/timeline-1.0/audio, operation concat | 1 to 20 parts, no re-synthesis |
| Place under video | POST /v1/timeline-1.0/render | audio.url or audio.parts[] as the spine |
Which engine does Sume's TTS use?
Sume's TTS 1.0 has no engine picker: model and model_id are rejected with 400. The separate TTS Router lists explicit models at GET /v1/tts-router/models. Check that catalog for what is routable today rather than assuming it matches an engine named in Google's recap.
What output format should I ask for?
The default is mp3 at 44100 Hz and 128 kbps. When the audio will feed an avatar mux, the OpenAPI notes say to pass wav with pcm_s16le at 44100 explicitly. For a plain video voiceover the default is fine.
Can Sume match Vids captions?
Sume has its own caption styles through POST /v1/video-captions; see the companion post on aligning captions to a script. It does not copy Vids' caption options.
Where does a scripted pipeline beat an editor?
An editor is quick for one video. A pipeline pays off at the tenth: the script is a file, the voice is a parameter, and the render is a job with an id. When a price or a claim changes, you edit one line and re-run, instead of reopening a project and re-recording.
It also helps review. A script in version control has a history, and each voiceover job records the transcript it was generated from. A reviewer reads the text, and the audio follows from it.
What does a minimal run look like?
Generate each paragraph as its own TTS job, keep the ids in order, join the audio with a concat job, and send the joined file to Timeline as the audio spine. If a paragraph reads badly, regenerate only that one and join again. The sentence segmentation option can return gapless per-sentence slices when you need to line visuals up with sentences.
Sources
Related posts
More in Comparisons
- gpt-4o-mini-tts instructions vs Sume TTS speed, volume, emotion
OpenAI steers gpt-4o-mini-tts with free-text instructions. Sume TTS exposes speed 0.6-1.5, volume 0.5-2 and a short emotion string. How the controls compare.
- GPT-6 Sol vs Luna vs Astra: context, prices, cutoffs
GPT-6 Astra, Sol and Luna share a 1,050,000-token window but differ 100x in price. A table from OpenAI's own pages, plus which one Sume runs Formats on.
- gpt-image-2.5 partial image streaming vs Sume's 400
OpenAI streams 0 to 3 partial images for gpt-image-2.5. Sume returns 400 streaming_not_supported for stream: true. What to send instead: a job and a poll.
- gpt-image-2.5 rate tiers (5 to 250 IPM) vs Sume plan queue
OpenAI limits gpt-image-2.5 Flare from 5 images per minute at Tier 1 to 250 at Tier 5. Sume limits processing concurrency by plan and queues the rest.
Written by Sume