Gemini Live Translate takes audio only: what to do with a script

Google's Live Translate model rejects text input. If your source is a written script, translate the text yourself and voice it with Sume TTS jobs.

5 min readSume
All posts

Gemini Live Translate cannot take a script. Google's Live Translate model for the Gemini Live API accepts audio input only, so you cannot send it a script, a caption file or a product description and get a translated voice back. If your source is written text, translate the text with a text translator and voice the result with a text-to-speech job; Sume runs that second half as separate, billable jobs you can inspect.

The constraint comes from Google's own page, Live translation with Gemini Live API, which says that only audio input is supported for translation and that text input is not. The model id on that page is gemini-3.5-live-translate-preview, and the page lists 70+ languages. This post is about what that audio-only rule means for someone who starts from a script.

What does the Live Translate page say it accepts?

Everything below is from the Google page, read on 2026-10-03. Anything the page does not state is left out here, including session length and price, which the page does not list.

Gemini Live Translate input and output, from Google's page (read 2026-10-03)
ItemWhat Google documents
Model idgemini-3.5-live-translate-preview
InputRaw 16-bit PCM, 16 kHz, mono, little-endian
OutputRaw 16-bit PCM, 24 kHz, mono, little-endian
Text inputNot supported, audio only
Target languagetargetLanguageCode, BCP-47, default en
TranscriptsinputAudioTranscription and outputAudioTranscription can be enabled

Why does an audio-only model matter for a script?

A script has no audio. To feed it to Live Translate you would first have to record or synthesize the source language, which adds a speech step whose only job is to be translated back out of speech. The result is a longer chain than the task needs: text to speech, speech to translated speech, and a voice that Google's page says can drift after long pauses or with several speakers.

The same page lists other limits worth reading before you plan around it: language detection can struggle with accents or similar languages, and background audio filtering can be inconsistent. None of these block a live call, where a person can repeat themselves. They matter more for a finished marketing clip, where nobody is there to correct the output.

What is the shorter path from script to translated voice on Sume?

Sume does not translate text on a route of its own; the docs list no translation endpoint. What it does ship is the voice and the media plumbing. A text-to-speech job records which voice, language and speed it used (model_id, voice, language, output_format, speed), as the jobs and results page describes, so the second language can reuse the same settings.

The path for a script is therefore: translate the text with the tool you trust, send each line to a text-to-speech job with the target language, then join the lines with timeline audio. A concat takes 1 to 20 parts, joins in the sample domain with no gap at the seams, and costs $0.01 flat per job according to that page.

When is Live Translate still the right tool?

When the source is a person speaking now. A meeting, a call or a stage talk is already audio, and the model's whole design is to stay a few seconds behind the speaker. Sume does not do live streaming translation; its audio jobs work on finished files.

For a recording you already have, the two services meet in the middle. Audio detach can pull the track from a Sume-hosted video as mono 16 kHz wav, which matches the sample rate Google lists for input, though the model expects raw PCM and not a wav container, so strip the header before streaming.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: detach-16k-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "format": "wav", "channels": "mono", "sample_rate": 16000}'

What should you check before choosing?

Start from the source, not the model. Ask three questions and the choice usually falls out.

  • Is the source live speech? Use a live model; Sume has no live route.
  • Is the source a script or a caption file? Translate the text, then voice it per language with TTS jobs.
  • Is the source a finished clip? Transcribe it, review the text, translate, voice, join. Each step is a job you can read before paying for the next.
  • Does the clip need lip-matched video? Sume's docs say video models do not lip-sync to generated TTS; talking faces use Fabric with an accepted still plus TTS.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume