Teams translated captions vanish after the meeting: keep them
Microsoft says Teams translated captions and transcripts are only available during the meeting. Caption the recording afterwards with Sume STT and burned cues.

Microsoft's support page says the translated live captions and transcripts from multilingual speech recognition are only available during the meeting and are not accessible after it ends. To keep a captioned copy, record the meeting, import the recording, transcribe it with Sume and burn captions onto a clip.
The details below come from Multilingual speech recognition in Microsoft Teams, read on 2026-10-03.
What does the Teams page say about languages and licences?
Teams lists nine languages for multilingual speech recognition: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese and Korean. Each participant picks their own spoken language from the language settings in the meeting controls, and live captions and transcripts then appear in each participant's selected language.
The feature needs a Microsoft Copilot licence or Teams Premium. If the organizer has one of them, the page says all participants can use translated captions and transcription without a licence of their own. The page also says the feature is not currently supported in Teams town halls.
| Item | What Microsoft's page says |
|---|---|
| Languages | English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean |
| Who chooses the language | Each participant, in meeting settings |
| After the meeting | Translations are not accessible |
| Licence | Copilot or Teams Premium for the organizer covers participants |
| Town halls | Not currently supported |
How do you get a durable transcript from the recording?
Import the recording first. Sume routes read only a workspace's own media.sume.com files, and the media inputs page covers POST /v1/media-imports. Then call video inspect with transcribe: true. The public rate is $0.01 per audio minute, and the duration hint maxes at 600 seconds, so a long meeting needs ranges or a shorter clip.
Pass language_code when you know the speech, for example en or ko; omit it and Sume auto-detects. Add segmentation.mode: "sentence" and the result includes gapless sentence segments[] shaped like caption lines.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: meeting-stt-001" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/meeting.mp4", "frames": false, "transcribe": true, "language_code": "en",
"duration_seconds": 600, "segmentation": {"mode": "sentence"}}'What does the transcript not give you?
It gives you the spoken language, not the translations each attendee saw. Teams made those on the fly and, per Microsoft, does not keep them. Sume's speech-to-text returns text in the language spoken, so a Portuguese-speaking attendee's translated view has to be rebuilt: translate the sentence segments with a translator you choose, then burn the result.
Sume's docs list no speaker labels for speech-to-text, so do not promise who-said-what output from it. If attribution matters, record that separately in the meeting tool before you rely on the recording.
How do you burn captions for the clips people will share?
Cut the moment you want to share, then send the translated sentences as cues to video captions. A cue is text, start and end in seconds, and cues skip speech-to-text. Each standalone job is $0.20 for videos up to 60 seconds under the current fixed estimate, so a clip in three languages is three jobs.
Check the clip has speech first. Without cues, a silent clip fails with caption_no_speech. With cues the job burns your copy at your times, even over silence.
- Record before the meeting starts; the live captions will not exist afterwards.
- Name one owner for approving the transcript before translation.
- Keep the original recording in the workspace so each new language reuses it.
Sources
Related posts
More in Use cases
- Template bulk edits on YouTube Shorts: what to vary per row
YouTube's Oct 1, 2026 originality update names template-based bulk changes as not original. How to make each row of a Sume bulk run differ in substance.
- Avatar reaction 3.49 vs 3.06 predicted: run your own two-clip test
HeyGen's survey says avatar users saw warmer reactions than skeptics predicted. Test your own audience with two Sume avatar clips, one script and safe retries.
- Editing a TikTok ad's video triggers a new review: plan for 24 hours
TikTok says saving or editing an ad's creative starts a new review, and most ads clear in about 24 hours. Plan a swapped AI video, and cut one with Video trim.
- TikTok ad static image rule: 50% of the video, and AI slideshows
TikTok's ad policy says still images should not fill more than 50% of a video ad and the ad needs clear audio. How to check an AI clip before you upload.
Written by Sume