Gemini Live voice agent tutorial: which pieces Sume covers
Gemini 3.8 Live went GA on 2026-09-15 for live voice agents. Sume does not host live sessions; it makes async TTS, STT and music jobs. What fits.

Build the live conversation with Google's Gemini Live API; Sume does not host live voice sessions. Sume's text-to-speech is an async job with a poll or webhook, documented as non-streaming, so it cannot be the speaking half of a live agent. It can produce the fixed audio around one: greeting and hold clips, test utterances, and transcripts of recordings.
Google's Gemini API changelog (read 2026-10-02) lists GA of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 2026-09-15, the second with background reasoning during audio. The tutorial for those models is on Google's pages; this post is about the work next to it.
What can Sume make for a voice agent?
Anything that is a file rather than a conversation. The table separates the two.
| Need | Gemini Live | Sume |
|---|---|---|
| Two-way spoken conversation | Yes, the Live API | No live sessions |
| Fixed greeting or hold message | Possible | TTS job returns an audio file |
| Phone-grade test audio | Not covered here | TTS output_format allows 8000 Hz and pcm_mulaw |
| Transcript of a recorded call | Not covered here | STT job with word timings |
How do I make phone-style test audio?
Sume's TTS output_format accepts a sample_rate of 8000 and an encoding of pcm_mulaw or pcm_alaw in a wav or raw container, so you can generate caller utterances in the format a phone leg uses and replay them against your agent.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caller-line-001" \
-d '{"transcript": "I need to move my appointment to Friday.", "avatar_handle": "@your_avatar", "output_format": {"container": "wav", "sample_rate": 8000, "encoding": "pcm_mulaw"}}'Does Sume answer the call?
No. Sume has no telephony or streaming endpoint. Wire the Gemini Live session to your own telephony layer, and use Sume files where a recorded line is acceptable.
When should I choose a job over a live session?
When no one waits on the audio. Narration, ad reads and fixed prompts are jobs. A conversation is a session.
What should I not expect from Sume here?
Sume has no streaming audio endpoint, no live session and no telephony. A caller cannot interrupt a Sume job, and a job cannot react to what the caller says. If your product is the conversation, build it on the live API and treat Sume as a file factory beside it.
Be careful with timing too. A TTS job returns after generation, and the synchronous wait is bounded at 30 seconds; it bounds the HTTP wait, not the job. For anything a caller hears in the first second, pre-render the file ahead of time and store its URL.
Where do the pre-rendered lines help most?
Fixed audio is where a job beats a session: the greeting, the hold message, the recording notice, the closing line, and error prompts. Generate them once, store the media.sume.com URL with the script text, and replay them. When the wording changes, regenerate that one file with a new idempotency key and update the URL.
What is a sensible split of work?
Build and test the conversation on Google's side first. Once the agent works, list every line it says that never changes, render those through Sume, and keep the live model for the parts that need judgment. That keeps the live session short and the fixed audio consistent from call to call.
Sources
Related posts
More in Comparisons
- Where is Gemini Omni 1.1 Flash? Google's named partners and Sume
Google names Adobe Firefly, Figma Weave, Runway and GMI Cloud as production users of Omni 1.1 Flash. What that means if you want one API with job ids.
- Gemini Omni Flash for client work: app, Flow, Shorts, or an API?
Gemini Omni Flash is in five places. Who each is for, what is free, and when a client project needs an API with job ids rather than an app.
- Gemini Omni won't take photos of minors in the EEA, UK or Switzerland
Google's Omni API guide blocks uploading images with minors in the EEA, UK and Switzerland. What the guide says, what Sume's docs say, and what to check first.
- Gladia Starter $0.61 per hour vs Sume STT $0.60 per hour
Gladia Starter prices async transcription at $0.61 an hour with 50 euros of free credit. Sume STT works out to $0.60 an hour. What differs besides the price.
Written by Sume