Is Gemini 3.8 Flash TTS in the Live API? No, per Google's model page

Google's model page marks the Live API unsupported for Gemini 3.8 Flash TTS. For live voice use Gemini 3.8 Live; Sume TTS is for async audio jobs.

4 min readSume
All posts

No. The Gemini 3.8 Flash TTS model page lists the Live API as not supported, along with function calling, structured outputs, thinking, code execution, file search, image generation and URL context. Google's changelog shows Gemini 3.8 Live as the separate model for live conversation.

What Google lists as supported

The same model page lists Batch API, Flex, Priority and caching. So the model is built for request and response speech generation, not a streaming two-way session. The changelog records TTS going GA on Sept 22 and Gemini 3.8 Live arriving on Sept 15.

Gemini 3.8 TTS support, read 2026-10-07
CapabilityFlash TTS
Live APINot supported
Function callingNot supported
Batch APISupported
Flex and PrioritySupported
CachingSupported

What this means for a voice agent

If a customer needs sub-second back-and-forth, build on the live model, not TTS. If you need a finished read of an ad script, TTS is the right shape. Mixing them (an agent that generates a take on demand) means two separate calls with two separate prices.

Where Sume sits

Sume TTS 1.0 is an asynchronous job: you post a transcript with an idempotency key and fetch the audio result (see jobs and results). It is not a live, interruptible voice session, and Sume does not claim to be one.

For a talking clip rather than a conversation, generate the voiceover first, then attach it to a face; that path is covered in our related post.

Sources

Related posts

More in Models

All Models posts

Written by Sume