Gemini Live voice agent tutorial: which pieces Sume covers

Gemini 3.8 Live went GA on 2026-09-15 for live voice agents. Sume does not host live sessions; it makes async TTS, STT and music jobs. What fits.

4 min readSume
All posts

Build the live conversation with Google's Gemini Live API; Sume does not host live voice sessions. Sume's text-to-speech is an async job with a poll or webhook, documented as non-streaming, so it cannot be the speaking half of a live agent. It can produce the fixed audio around one: greeting and hold clips, test utterances, and transcripts of recordings.

Google's Gemini API changelog (read 2026-10-02) lists GA of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 2026-09-15, the second with background reasoning during audio. The tutorial for those models is on Google's pages; this post is about the work next to it.

What can Sume make for a voice agent?

Anything that is a file rather than a conversation. The table separates the two.

Live agent versus Sume jobs. Google changelog and Sume docs, read 2026-10-02.
NeedGemini LiveSume
Two-way spoken conversationYes, the Live APINo live sessions
Fixed greeting or hold messagePossibleTTS job returns an audio file
Phone-grade test audioNot covered hereTTS output_format allows 8000 Hz and pcm_mulaw
Transcript of a recorded callNot covered hereSTT job with word timings

How do I make phone-style test audio?

Sume's TTS output_format accepts a sample_rate of 8000 and an encoding of pcm_mulaw or pcm_alaw in a wav or raw container, so you can generate caller utterances in the format a phone leg uses and replay them against your agent.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caller-line-001" \
  -d '{"transcript": "I need to move my appointment to Friday.", "avatar_handle": "@your_avatar", "output_format": {"container": "wav", "sample_rate": 8000, "encoding": "pcm_mulaw"}}'

Does Sume answer the call?

No. Sume has no telephony or streaming endpoint. Wire the Gemini Live session to your own telephony layer, and use Sume files where a recorded line is acceptable.

When should I choose a job over a live session?

When no one waits on the audio. Narration, ad reads and fixed prompts are jobs. A conversation is a session.

What should I not expect from Sume here?

Sume has no streaming audio endpoint, no live session and no telephony. A caller cannot interrupt a Sume job, and a job cannot react to what the caller says. If your product is the conversation, build it on the live API and treat Sume as a file factory beside it.

Be careful with timing too. A TTS job returns after generation, and the synchronous wait is bounded at 30 seconds; it bounds the HTTP wait, not the job. For anything a caller hears in the first second, pre-render the file ahead of time and store its URL.

Where do the pre-rendered lines help most?

Fixed audio is where a job beats a session: the greeting, the hold message, the recording notice, the closing line, and error prompts. Generate them once, store the media.sume.com URL with the script text, and replay them. When the wording changes, regenerate that one file with a new idempotency key and update the URL.

What is a sensible split of work?

Build and test the conversation on Google's side first. Once the agent works, list every line it says that never changes, render those through Sume, and keep the live model for the parts that need judgment. That keeps the live session short and the fixed audio consistent from call to call.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume