Gemini 3.8 Live or TTS plus Fabric for a talking clip on Sume
Google shipped speech-to-speech and TTS models in September. A talking video on Sume is a TTS voice plus a Fabric clip. Why the two are different products.

If Google's new voice models talk, can they make a talking video? Gemini 3.8 Live is a conversation model; a talking video needs audio, a face and lip sync. On Sume the talking shot is built from a TTS voice and a Fabric clip, not from a speech-to-speech session.
Google's September 2026 update (read 2026-10-04) says it released Gemini 3.8 Live for everyday conversation and Gemini 3.8 Live Extended Thinking, described as the top-rated speech-to-speech model for multi-step reasoning without lag, both available to developers in Google AI Studio and the Gemini API. It also lists Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS for custom voices and scene dialogue.
Two different jobs
A speech-to-speech model listens and answers in audio in real time. That is an assistant. A talking clip is an asset you render once and reuse. The models overview states Sume's rule plainly: every on-camera speaking shot is Fabric with an accepted still plus TTS, and video models do not lip-sync to generated TTS or to a later voice-over.
Side by side
The Google column is from Google's post; the Sume column is from Sume's docs.
| Question | Gemini 3.8 Live (Google) | TTS plus Fabric (Sume) |
|---|---|---|
| Output | Spoken replies in a live session | A finished video file with a face |
| Where it runs | Google AI Studio, Gemini API, Gemini apps | Sume API jobs, polled or via webhook |
| Face and lip sync | Not part of the announcement | Fabric clip from a still and audio |
| Wordless shots | Not applicable | Auto image, then inspect, then Auto video |
The Sume recipe
Make the voice with a TTS job, render a posed still, then call POST /v1/veed/fabric-1.0 with audio_url, a measured duration_seconds and exactly one visual source: image_url of the inspected still, or avatar_handle when the user named that avatar. The two visual fields are mutually exclusive. MiniMax H3 Max Lip Sync takes the same still and audio body with audio of 5 to 14.8 seconds.
If you only want a presenter and a script, Avatar video takes a script and an avatar handle, 4 to 60 seconds, and handles the voice for you.
Use Google's live models for a talking assistant. Use Sume when the result has to be a video you can caption, trim and publish.
Sources
Related posts
More in Comparisons
- Gemini API webhooks for batch jobs: keep a poll backup, as on Sume
Gemini API release notes say webhooks replace polling for Batch and long-running operations. What changes in a client, and why Sume keeps a polling backup.
- Gemini batch inline 20 MB or 2 GB file vs Sume's 100 inline items
Gemini Batch takes inline requests under 20 MB or a file up to 2 GB. Sume bulk takes 1 to 100 inline items. How to size a holiday video batch for each.
- Gemini batch job EXPIRED after 48 hours vs Sume run expires_at
A Gemini batch job ends EXPIRED after 48 hours pending or running. A Sume Format run is finalized as failed at expires_at. Compare the two clocks.
- Gemini Omni Flash 1,240 vs Seedance 2.0 1,225: what 15 Elo means
Hedra lists Gemini Omni Flash first at 1,240 Elo and Seedance 2.0 4K second at 1,225. A 15-point gap is about a 52% win rate; check limits before choosing.
Written by Sume