Gemini Live Translate transcripts as subtitles: you supply timing

Google's Live Translate can return input and output transcripts, but the page lists no word times. To burn subtitles, time each line, then send cues to Sume.

5 min readSume
All posts

Yes, you can get text out of Gemini Live Translate and burn it as subtitles, but the text arrives without the timing a caption needs, so you have to supply it. Google's page says you can enable inputAudioTranscription and outputAudioTranscription to receive transcripts of the input and the translated audio; it does not describe word or sentence times. Sume's caption cues take text, start and end, so the missing piece is the clock.

The Google facts here come from Live translation with Gemini Live API, read on 2026-10-03.

What does the Google page give you?

The page lists the model id gemini-3.5-live-translate-preview, 16 kHz mono PCM in, 24 kHz mono PCM out, a targetLanguageCode in BCP-47 form, and the two transcription switches. It says audio only is supported as input. It does not list pricing or session duration, so neither is stated here.

Because the model streams a few seconds behind the speaker, any timestamp you attach when a transcript fragment arrives is late by that lag. A caption stamped on arrival would trail the voice.

What a caption job needs against what the Google page lists (read 2026-10-03)
Caption needGemini Live Translate pageSume video captions
Text of translated speechoutputAudioTranscriptioncues[].text
Start and end secondsNot describedcues[].start and cues[].end
Source transcriptinputAudioTranscriptionOptional; STT runs when cues are omitted
Target languagetargetLanguageCodelanguage is only an STT hint

How do you get timing that matches the video?

Use a second source for the clock. Run the finished recording through video inspect with transcribe: true and segmentation.mode: "sentence". The result includes gapless sentence segments[] with start and end times, and word-level timings as words[]. Those times belong to the original speech, not to the delayed translation.

Then pair sentences. If the translated transcript has the same number of sentences as the source, map them one to one and reuse the source start and end. If the counts differ, have a person align the lines before burning. Misaligned captions are worse than none.

How do you burn the paired lines?

Send the paired lines as cues to video captions. The page says cues (or segments) skip speech-to-text and burn exactly your copy at your times, so the translation is never re-transcribed. A standalone job is $0.20 for videos up to 60 seconds under the current fixed estimate, and each language is its own job.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: live-translate-cues-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/clean.mp4", "style": "slam",
       "cues": [{"text": "Welcome to the demo.", "start": 0.4, "end": 2.1},
                {"text": "Today we cover pricing.", "start": 2.3, "end": 4.0}]}'

What can go wrong?

Google's page lists several limits that affect text quality: voice consistency issues after long pauses or with several speakers, trouble detecting language with accents or similar languages, and inconsistent background audio filtering. If language detection slips, the transcript for that stretch can be in the wrong language.

Check the translated text before burning it. Sume's job is to put your words on frames at your times; it will not catch a mistranslation.

  • Read every cue once for names, numbers and product terms.
  • Keep each cue short; a line that does not fit the screen is a design problem, not a timing one.
  • Pass the right style for the script; Korean text on a Latin style returns a 400.
  • Keep the source segments so a corrected line can be re-burned without new timing.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume