Teams Interpreter: 9 languages, English pivot vs Sume steps

Microsoft says Teams Interpreter turns speech into English text, translates, then speaks. Sume runs that chain as separate jobs so you can check the text.

5 min readSume
All posts

How does the Teams Interpreter agent translate speech?

Microsoft Learn describes Interpreter as real-time speech-to-speech translation built on Azure services in four steps: speech recognition converts spoken language into English text, machine translation turns that English text into the selected languages, text-to-speech produces the translated speech, and a bot carries meeting audio to the cloud and back. Participants can have Interpreter simulate their own voice or pick a preset voice (Ava, Andrew or Fable Turbo; Ava is the default).

That is the same shape as any dubbing pipeline: recognize, translate, speak. What differs is whether you can look at the text in the middle. In a live meeting you cannot. In a batch pipeline you can, and that is where most dubbing errors are cheapest to catch.

Which languages and licences does it need?

The Learn page lists nine languages for speaking and listening: Chinese (Mandarin), English, French, German, Italian, Japanese, Korean, Portuguese and Spanish. It requires a Microsoft 365 Copilot licence, includes 20 hours of interpretation per user per month, and says access beyond that depends on available capacity. It is supported in scheduled and channel meetings, 1:1 Teams-to-Teams VoIP calls and Teams Rooms on Windows, but not in Teams events, 1:1 PSTN calls or free Teams.

Plan for the exclusions early: a webinar run as a Teams event, or a call that includes a phone participant on PSTN, will not get Interpreter, and the page offers no workaround. Admins control it with the -AIInterpreter and -VoiceSimulationInInterpreter parameters of Teams policy cmdlets. The page says voice samples and biometric data are not stored when voice simulation is on.

What does an English pivot mean for quality?

The Learn page says recognition produces English text, and machine translation turns English into the chosen languages. For a Korean-to-Japanese conversation, that implies the meaning passes through English on the way. We can only read what the page says; it does not describe how it handles language pairs internally beyond those four steps.

For your own pipeline, the lesson is practical: a pivot language is a risk to check, not a given. Names, honorifics and idioms can flatten when two translations are chained. On Sume you choose how translation happens, because it is not a Sume step. Translate directly from the source language with whatever tool or agent you trust, and keep the source transcript beside the translation so a reviewer can compare them.

How do you build a reviewable version on Sume?

Start from the meeting or lesson recording as a hosted video, then run the chain as separate jobs. Detach the audio, transcribe with sentence segmentation, translate and review the sentences, speak each approved line with TTS and join the audio. Every step returns a result you can inspect before the next one is paid for.

Details that save time: video inspect can transcribe a clip directly when transcribe: true is set (probe and stills are unbilled); STT takes at most 10 minutes per job; TTS needs language set for every non-English transcript or it defaults to English; and a voice whose language does not match gets a warning you confirm explicitly. The mismatch warning post explains the retry.

Teams Interpreter and a Sume pipeline, read 2026-10-02
QuestionTeams Interpreter (Microsoft Learn)Sume pipeline
TimingReal time in a meetingBatch jobs on a recording
Languages9 listedPer step; no single list
Text visible before speechNoYes
VoiceSimulated own voice or 3 presetsTTS voice you select
Included usage20 hours per Copilot user per monthPay per job

When is each right?

Use Teams Interpreter when colleagues need to understand each other now and your organisation already pays for Copilot. Use a Sume pipeline when you need an artifact: a translated training video, a recap clip, a captioned announcement. One practical rule: never publish a live translation as the official record without a person reading it first. Neither Microsoft's page nor Sume makes any accuracy guarantee, and a one-line error in a safety or compliance video is the kind of mistake you cannot patch after it has circulated.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume