Google Meet speech translation: caption the recording with Sume

Meet translates live speech between English and five languages, one pair per call. Google's pages do not say recordings carry it, so caption the file yourself.

5 min readSume
All posts

What does Google Meet's speech translation do, and what does it leave out?

Meet's speech translation lets people speak their own language while others hear an audio translation. Google's February 2026 post says it covers bidirectional translation between English and Spanish, French, German, Portuguese and Italian, and that it dubs audio over the original speech to mimic the speaker's tone and cadence rather than showing captions. Only one language pair can be active in a meeting at a time.

The two Google pages we read do not say whether a saved recording captures the translated audio, translated captions, or both. If you need a translated, shareable file after the call, do not assume the recording has it. Build it from the recording.

Which plans and devices does it cover?

Per the same pages, it is offered on Business Standard and Plus, Enterprise Standard and Plus, Frontline Plus, Google AI Pro and Ultra, plus an AI Ultra for Business add-on and an AI Pro for Education add-on. The April 8, 2026 post covers the mobile rollout to Android and iOS, with Rapid Release domains starting April 8 and Scheduled Release domains April 23. Conference-room hardware users can hear translations, but their own speech is not translated.

Two limits to plan around: one language pair per meeting, and the English-centred pair list. A call that mixes Spanish and German speakers cannot use it for both at once.

How do you caption a meeting recording in two languages with Sume?

Sume does not join calls and has no streaming. It works on a recording you upload as a file, in short jobs. The chain is: detach the audio, transcribe it, translate the text, then burn the approved lines as captions.

Speech-to-text takes at most 10 minutes per job, and audio detach caps the source at 1,800 seconds and each output at 900 seconds, so a 45-minute meeting needs several ranges. Each range returns a new audio_url you pass to STT. Request mono 16 kHz wav, the speech-to-text shape. The recording must already be a video on this workspace's media.sume.com host; import it first with POST /v1/media-imports, because audio detach does not fetch arbitrary URLs. A source with no audio track fails with detach_source_has_no_audio, which is worth checking when a recording is screen-only.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: meet-recording-part-1" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/meeting.mp4",
    "format": "wav",
    "channels": "mono",
    "sample_rate": 16000,
    "range": { "start": 0, "end": 600 }
  }'

What do you do with the transcript?

Transcribe each range with a language_code hint (en, es and so on) and segmentation: {"mode": "sentence"}, then offset each part's timestamps by its range start so the times line up with the whole recording. Translate the sentences, review them, and send them to video captions as cues with text, start and end.

Standalone caption jobs take a public HTTPS video_url and are priced at $0.20 per job for videos up to 60 seconds, so burning a long meeting means trimming it into clips first. For a one-hour call that is rarely the right output; it suits a highlight, a decision clip or a customer-facing recap. See meeting minutes from a recording for the text-only version.

Live translation vs a processed recording, read 2026-10-02
NeedMeet speech translationSume on a recording
When it happensDuring the callAfter the call
OutputDubbed audio for listenersCaptioned clips or text
LanguagesEnglish with 5 languages, one pairYour translation; STT language hint per range
Shareable translated fileNot stated by Google's pagesYes, a burned-in MP4 per language
StreamingYesNo, batch jobs

Should you rely on live translation for compliance records?

No. A live AI translation is a convenience for participants. If a translated record matters (training, legal, regulated customer calls), keep the original recording, keep the transcript and have a person review the translated text before it is published or archived. The Google pages do not make claims about accuracy guarantees, and neither does Sume.

If the call mixed languages, tell the transcription step what to expect: Sume's language_code is only a hint, and omitting it means automatic detection, which is the safer choice for a segment where two people switch languages. Run a short test range before you process the whole file.

A practical split: use Meet translation to run the conversation, then use Sume to turn the best five minutes into a captioned clip for each market you serve.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume