Gemini 3.5 Live Translate (70+ languages) vs Sume batch dubbing

Google's Gemini 3.5 Live Translate runs a few seconds behind a live speaker. For a finished clip you can translate in steps and review before any audio is made.

5 min readSume
All posts

What is Gemini 3.5 Live Translate, and is it a dubbing tool?

Gemini 3.5 Live Translate is a speech-to-speech translation model from Google, announced June 9, 2026. Google's post says it supports 70+ languages with automatic detection, keeps the speaker's intonation, pacing and pitch, and generates speech continuously so the translation stays just a few seconds behind the speaker. It is built for live conversation, not for producing a polished localized video.

That distinction matters for anyone searching for a way to dub a finished clip. A live model optimizes for staying in sync with a person who is still talking. A finished clip has the opposite shape: you know the whole script in advance, so you can translate sentence by sentence, check the lengths and only then make audio.

Where is it available and what are the limits?

The post lists three places. The Google Translate app is rolling out globally on Android and iOS, including a listening mode on Android. The Gemini Live API and Google AI Studio carry it in public preview for developers. Google Meet had it in private preview for select Workspace customers, with broader rollout planned later in 2026. The post says it enables 2,000+ language combinations in a single meeting.

Google also states that all audio output is watermarked with SynthID so it can be detected. If you plan to republish translated audio, assume the watermark travels with the file and tell viewers the audio is synthetic where a platform requires it.

Why a finished clip is better translated in steps

A live model cannot go back. If it mishears a product name in minute two, that error is already spoken. A step pipeline lets you fix the text before any voice exists, and lets you retry just the sentence that came out too long.

On Sume the pieces are separate jobs. Audio detach gives you the clip's audio, STT returns the words and sentence segments with timings, you translate and review the sentences, TTS speaks each one with language set to the target, and timeline audio joins up to 20 parts with no re-synthesis and no silence at the seams. The sentence-segment technique is covered in keeping a translated line on its slot.

  • Nothing is spoken until the text is approved.
  • A bad line is a one-sentence retry, not a rerun of the whole clip.
  • Every intermediate file stays on media.sume.com for audit.

What does Sume not do that a live model does?

Sume has no streaming session. STT, TTS and caption jobs are asynchronous: you submit, and poll or receive a webhook. The default wait for a synchronous call is bounded at 30 seconds, and the docs tell you to use async plus a callback for anything longer. There is no way to pipe a microphone through Sume and hear a translation a few seconds later.

It also does not detect language in a way that mixes several in one stream for you. The STT language_code is an optional BCP-47 hint; omit it and the job auto-detects. For a speaker who switches mid-sentence, review the transcript instead of assuming the detector caught every switch.

Live translation vs batch dubbing, read 2026-10-02
QuestionGemini 3.5 Live TranslateSume steps
Latency modelA few seconds behind the speakerAsync jobs, minutes
Languages70+ with automatic detectionPer step; set language on TTS
Review before audioNo, it speaks as it goesYes, between jobs
OutputSpoken translation (SynthID-watermarked)Files you keep
Best forConversations, meetings, travelAds, shorts, training clips

Which should you use?

Use Google's live translation when two people need to talk now. Use a stepwise pipeline when the result is a video that will be watched thousands of times, because one wrong word is repeated every time. For a mixed workflow, run the live call with Meet or the Translate app, then bring the recording through Sume to make the polished version for each market. The real-time vs async TTS post compares the two job shapes in more detail.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume