Gemini 3.8 Live or TTS plus Fabric for a talking clip on Sume

Google shipped speech-to-speech and TTS models in September. A talking video on Sume is a TTS voice plus a Fabric clip. Why the two are different products.

6 min readSume
All posts

If Google's new voice models talk, can they make a talking video? Gemini 3.8 Live is a conversation model; a talking video needs audio, a face and lip sync. On Sume the talking shot is built from a TTS voice and a Fabric clip, not from a speech-to-speech session.

Google's September 2026 update (read 2026-10-04) says it released Gemini 3.8 Live for everyday conversation and Gemini 3.8 Live Extended Thinking, described as the top-rated speech-to-speech model for multi-step reasoning without lag, both available to developers in Google AI Studio and the Gemini API. It also lists Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS for custom voices and scene dialogue.

Two different jobs

A speech-to-speech model listens and answers in audio in real time. That is an assistant. A talking clip is an asset you render once and reuse. The models overview states Sume's rule plainly: every on-camera speaking shot is Fabric with an accepted still plus TTS, and video models do not lip-sync to generated TTS or to a later voice-over.

Side by side

The Google column is from Google's post; the Sume column is from Sume's docs.

Live voice model versus rendered talking clip, read 2026-10-04
QuestionGemini 3.8 Live (Google)TTS plus Fabric (Sume)
OutputSpoken replies in a live sessionA finished video file with a face
Where it runsGoogle AI Studio, Gemini API, Gemini appsSume API jobs, polled or via webhook
Face and lip syncNot part of the announcementFabric clip from a still and audio
Wordless shotsNot applicableAuto image, then inspect, then Auto video

The Sume recipe

Make the voice with a TTS job, render a posed still, then call POST /v1/veed/fabric-1.0 with audio_url, a measured duration_seconds and exactly one visual source: image_url of the inspected still, or avatar_handle when the user named that avatar. The two visual fields are mutually exclusive. MiniMax H3 Max Lip Sync takes the same still and audio body with audio of 5 to 14.8 seconds.

If you only want a presenter and a script, Avatar video takes a script and an avatar handle, 4 to 60 seconds, and handles the voice for you.

Use Google's live models for a talking assistant. Use Sume when the result has to be a video you can caption, trim and publish.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume