DeepL Voice keeps speakers' voices in 14 languages: Sume dubs
DeepL's September 15 release keeps each speaker's voice across 14 languages in live talk. A Sume dub picks a TTS voice id per line and does not clone.

DeepL's September 15, 2026 release makes voice-to-voice translation keep each speaker's voice distinct in 14 languages, in real time, inside Zoom, Microsoft Teams and Google Meet. A Sume dub works differently: you pick a text-to-speech voice per line, so a dubbed clip has a chosen voice, not a copy of the original speaker's.
That difference decides which tool you want. DeepL is built for live conversation. Sume is built for finished files, where every step is a separate job you can check. The facts about DeepL below come from its press release, read on 2026-10-03.
What did DeepL announce on September 15?
The release is headed "DeepL Voice now preserves your voice in real-time multilingual conversations". It describes voice preservation across 14 supported languages for voice-to-voice translation, with 30+ languages across the platform, and says natural expression such as hesitation and emotion carries through at low latency.
It also describes a desktop app for Windows and Mac that works with Zoom, Microsoft Teams and Google Meet, and says developers can reach DeepL Voice through its API. The release says the product powered real-time translation across 50 stages at Salesforce Dreamforce 2026 and was ranked first in Slator's assessment of AI-translated captions. The release does not give prices, so none appear here.
| Question | DeepL Voice release | Sume |
|---|---|---|
| Voice of the original speaker kept | Yes, 14 languages | No clone route in the docs; one TTS voice id per line |
| Real time | Yes, low latency | No; jobs run on finished files |
| Inside Zoom, Teams, Meet | Desktop app | Not applicable; you upload a recording |
| Output you can read before audio | Not stated | Transcript text between steps |
| Burned-in captions | Not stated | video-captions, $0.20 per job up to 60 seconds |
How does a Sume dub choose a voice?
A text-to-speech job records how its audio was made. The jobs and results page lists model_id, voice as { "mode": "id", "id": "..." }, language, output_format, speed and generation_config. Each is null when the request did not send it.
So a dub in a second language reuses a voice id and a language code per line, and you read the settings off the first job to make the next one sound the same. That keeps one consistent narrator across a series. It does not reproduce a particular person's timbre, and this post will not claim it does.
What does the Sume chain look like for a finished video?
Sume has no single dubbing call. You detach the audio, transcribe it, translate the text with a translator you pick, voice each line, and join the lines. Audio detach and timeline audio are each $0.01 per job, per their pages, and speech-to-text is $0.01 per audio minute per the video inspect page. Text-to-speech and translation are priced separately, so check GET /v1/catalog before you budget.
Because each step returns a file or text, you can stop after the transcript and have a person approve it. That is the trade against a live voice product: slower, but checkable.
- Live meeting with several languages: use a live tool such as DeepL Voice.
- Recorded webinar in one language going to five: run the chain per language and caption each.
- Brand narrator across 20 clips: one voice id, one language code per run.
- A founder's own voice in German: Sume's docs do not offer that, so use a tool that documents voice preservation.
Should you caption instead of dubbing?
For many recordings the cheaper answer is subtitles. A caption job takes a public HTTPS video URL, burns captions with a style and an optional language hint, and the standalone job costs $0.20 for videos up to 60 seconds under the current fixed estimate. Dubbing adds voice work for each language; captions do not.
If the audience will watch on mute, captions in the viewer's language cover them without any voice at all. If they will listen, dub. Either way, keep the original audio file so the next language does not start from scratch.
Sources
Related posts
More in Comparisons
- Does Sume have Veo 3.1? No. Which video models it lists instead
Sume's video catalog has no Veo model. See the models it does list, with duration and resolution ranges, and how Google's own docs now steer to Omni Flash.
- Eleven v4: 90+ languages or 99? ElevenLabs' own pages differ
ElevenLabs' launch post says more than 90 languages for Eleven v4; its docs page lists 99. How to plan around the gap, and what Sume lists instead.
- Eleven v4 drops the native accent; Sume keeps a language tag per voice
ElevenLabs' docs say v4 gives fluent target-language speech, not a preserved accent. Sume tags each voice with a language and checks for mismatch.
- Eleven v4 voice clone: 10 seconds or 1-2 minutes? Pages differ
ElevenLabs' launch post says an instant clone needs 10 seconds of audio; its docs page says one to two minutes. What Sume's Voices clone asks for instead.
Written by Sume