Teams Interpreter: 9 languages, English pivot vs Sume steps
Microsoft says Teams Interpreter turns speech into English text, translates, then speaks. Sume runs that chain as separate jobs so you can check the text.

How does the Teams Interpreter agent translate speech?
Microsoft Learn describes Interpreter as real-time speech-to-speech translation built on Azure services in four steps: speech recognition converts spoken language into English text, machine translation turns that English text into the selected languages, text-to-speech produces the translated speech, and a bot carries meeting audio to the cloud and back. Participants can have Interpreter simulate their own voice or pick a preset voice (Ava, Andrew or Fable Turbo; Ava is the default).
That is the same shape as any dubbing pipeline: recognize, translate, speak. What differs is whether you can look at the text in the middle. In a live meeting you cannot. In a batch pipeline you can, and that is where most dubbing errors are cheapest to catch.
Which languages and licences does it need?
The Learn page lists nine languages for speaking and listening: Chinese (Mandarin), English, French, German, Italian, Japanese, Korean, Portuguese and Spanish. It requires a Microsoft 365 Copilot licence, includes 20 hours of interpretation per user per month, and says access beyond that depends on available capacity. It is supported in scheduled and channel meetings, 1:1 Teams-to-Teams VoIP calls and Teams Rooms on Windows, but not in Teams events, 1:1 PSTN calls or free Teams.
Plan for the exclusions early: a webinar run as a Teams event, or a call that includes a phone participant on PSTN, will not get Interpreter, and the page offers no workaround. Admins control it with the -AIInterpreter and -VoiceSimulationInInterpreter parameters of Teams policy cmdlets. The page says voice samples and biometric data are not stored when voice simulation is on.
What does an English pivot mean for quality?
The Learn page says recognition produces English text, and machine translation turns English into the chosen languages. For a Korean-to-Japanese conversation, that implies the meaning passes through English on the way. We can only read what the page says; it does not describe how it handles language pairs internally beyond those four steps.
For your own pipeline, the lesson is practical: a pivot language is a risk to check, not a given. Names, honorifics and idioms can flatten when two translations are chained. On Sume you choose how translation happens, because it is not a Sume step. Translate directly from the source language with whatever tool or agent you trust, and keep the source transcript beside the translation so a reviewer can compare them.
How do you build a reviewable version on Sume?
Start from the meeting or lesson recording as a hosted video, then run the chain as separate jobs. Detach the audio, transcribe with sentence segmentation, translate and review the sentences, speak each approved line with TTS and join the audio. Every step returns a result you can inspect before the next one is paid for.
Details that save time: video inspect can transcribe a clip directly when transcribe: true is set (probe and stills are unbilled); STT takes at most 10 minutes per job; TTS needs language set for every non-English transcript or it defaults to English; and a voice whose language does not match gets a warning you confirm explicitly. The mismatch warning post explains the retry.
| Question | Teams Interpreter (Microsoft Learn) | Sume pipeline |
|---|---|---|
| Timing | Real time in a meeting | Batch jobs on a recording |
| Languages | 9 listed | Per step; no single list |
| Text visible before speech | No | Yes |
| Voice | Simulated own voice or 3 presets | TTS voice you select |
| Included usage | 20 hours per Copilot user per month | Pay per job |
When is each right?
Use Teams Interpreter when colleagues need to understand each other now and your organisation already pays for Copilot. Use a Sume pipeline when you need an artifact: a translated training video, a recap clip, a captioned announcement. One practical rule: never publish a live translation as the official record without a person reading it first. Neither Microsoft's page nor Sume makes any accuracy guarantee, and a one-line error in a safety or compliance video is the kind of mistake you cannot patch after it has circulated.
Sources
Related posts
More in Comparisons
- TikTok Smart Split vs Sume: turn a long video into vertical clips
TikTok Studio Web's Smart Split clips, reframes and captions long videos. What the same job takes with Sume trim, crop and captions, and what it does not do.
- TikTok product avatars skip shoes, hats, sunglasses and bracelets
TikTok's Symphony help page lists products its avatars cannot show. What to try for those SKUs, and what Sume's avatar product_image does and does not claim.
- Seedance 2.5 in Symphony takes 50 references: what Sume lists
TikTok says Seedance 2.5 in Symphony takes up to 50 image, video and audio references. Sume's docs give per-model reference types, so check the catalog first.
- Together AI dynamic rate limits (429, 503) vs Sume rate_limited
Together AI publishes no fixed tiers: limits track live capacity and your recent traffic. How its 429 and 503 map to Sume's rate_limited and queue_full.
Written by Sume