Gemini 3.5 Live Translate (70+ languages) vs Sume batch dubbing
Google's Gemini 3.5 Live Translate runs a few seconds behind a live speaker. For a finished clip you can translate in steps and review before any audio is made.

What is Gemini 3.5 Live Translate, and is it a dubbing tool?
Gemini 3.5 Live Translate is a speech-to-speech translation model from Google, announced June 9, 2026. Google's post says it supports 70+ languages with automatic detection, keeps the speaker's intonation, pacing and pitch, and generates speech continuously so the translation stays just a few seconds behind the speaker. It is built for live conversation, not for producing a polished localized video.
That distinction matters for anyone searching for a way to dub a finished clip. A live model optimizes for staying in sync with a person who is still talking. A finished clip has the opposite shape: you know the whole script in advance, so you can translate sentence by sentence, check the lengths and only then make audio.
Where is it available and what are the limits?
The post lists three places. The Google Translate app is rolling out globally on Android and iOS, including a listening mode on Android. The Gemini Live API and Google AI Studio carry it in public preview for developers. Google Meet had it in private preview for select Workspace customers, with broader rollout planned later in 2026. The post says it enables 2,000+ language combinations in a single meeting.
Google also states that all audio output is watermarked with SynthID so it can be detected. If you plan to republish translated audio, assume the watermark travels with the file and tell viewers the audio is synthetic where a platform requires it.
Why a finished clip is better translated in steps
A live model cannot go back. If it mishears a product name in minute two, that error is already spoken. A step pipeline lets you fix the text before any voice exists, and lets you retry just the sentence that came out too long.
On Sume the pieces are separate jobs. Audio detach gives you the clip's audio, STT returns the words and sentence segments with timings, you translate and review the sentences, TTS speaks each one with language set to the target, and timeline audio joins up to 20 parts with no re-synthesis and no silence at the seams. The sentence-segment technique is covered in keeping a translated line on its slot.
- Nothing is spoken until the text is approved.
- A bad line is a one-sentence retry, not a rerun of the whole clip.
- Every intermediate file stays on media.sume.com for audit.
What does Sume not do that a live model does?
Sume has no streaming session. STT, TTS and caption jobs are asynchronous: you submit, and poll or receive a webhook. The default wait for a synchronous call is bounded at 30 seconds, and the docs tell you to use async plus a callback for anything longer. There is no way to pipe a microphone through Sume and hear a translation a few seconds later.
It also does not detect language in a way that mixes several in one stream for you. The STT language_code is an optional BCP-47 hint; omit it and the job auto-detects. For a speaker who switches mid-sentence, review the transcript instead of assuming the detector caught every switch.
| Question | Gemini 3.5 Live Translate | Sume steps |
|---|---|---|
| Latency model | A few seconds behind the speaker | Async jobs, minutes |
| Languages | 70+ with automatic detection | Per step; set language on TTS |
| Review before audio | No, it speaks as it goes | Yes, between jobs |
| Output | Spoken translation (SynthID-watermarked) | Files you keep |
| Best for | Conversations, meetings, travel | Ads, shorts, training clips |
Which should you use?
Use Google's live translation when two people need to talk now. Use a stepwise pipeline when the result is a video that will be watched thousands of times, because one wrong word is repeated every time. For a mixed workflow, run the live call with Meet or the Translate app, then bring the recording through Sume to make the polished version for each market. The real-time vs async TTS post compares the two job shapes in more detail.
Sources
Related posts
More in Comparisons
- Gemini batch create is not idempotent: two jobs vs Sume bulk replay
Gemini's docs say sending the same batch creation request twice creates two batch jobs. Sume bulk runs replay the old queue for the same Idempotency-Key.
- Gemini Omni Flash for client work: app, Flow, Shorts, or an API?
Gemini Omni Flash is in five places. Who each is for, what is free, and when a client project needs an API with job ids rather than an app.
- Gladia Starter $0.61 per hour vs Sume STT $0.60 per hour
Gladia Starter prices async transcription at $0.61 an hour with 50 euros of free credit. Sume STT works out to $0.60 an hour. What differs besides the price.
- Google Chirp 3 Instant Custom Voice: 10 s consent, allowlist only
Google's Chirp 3 Instant Custom Voice is allowlisted, wants a 10-second consent clip, and keeps the cloning key on your side. How that compares to a Sume clone.
Written by Sume