DeepL Voice live subtitles vs translated captions on Sume
DeepL Voice shows live translated subtitles in Teams, Zoom and Meet. For a recorded clip, Sume burns the subtitles you author as cues, in Latin or Hangul text.

What does DeepL Voice offer, and where does Sume fit?
DeepL's product page describes DeepL Voice as live translation for meetings with translated subtitles you can follow across Microsoft Teams, Zoom Meetings and Google Meet, voice-to-voice translation that preserves how each speaker sounds, and an API for voice channels such as customer service and sales. It lists 40+ languages including English, French, German, Spanish, Japanese and Chinese, and states ISO/IEC 27001:2022 and SOC 2 Type 2 certification, GDPR and HIPAA compliance, and that data is not used for model training. Pricing is on a separate page that this post does not quote.
Sume covers a different moment: after the meeting or shoot, when you have a recorded clip and want subtitles burned into it. It does not translate live speech and has no meeting integrations.
What can Sume actually do with translated subtitles?
You provide the translated text; Sume renders it. A video captions job takes a public HTTPS video URL plus cues (or segments), each with text, start and end in seconds. Authored cues skip speech-to-text entirely, so the burned wording is exactly what you approved, and a silent clip works because no speech is needed.
If you would rather start from the audio, omit the cues: Sume transcribes with speech-to-text and burns that wording, with an optional script_text that aligns your script to the detected word timings. Alignment can fail with script_alignment_mismatch or script_alignment_failed, in which case simplify the script or omit it. SRT uploads and provider task ids are not accepted; pass phrase-level text as cues instead.
Does this replace a translation engine?
No. Sume has no translation endpoint in the API documentation we checked, so translation comes from your own tool, a translation service such as DeepL, or an agent step. The pipeline is: transcribe with Sume STT (or use your own transcript), translate elsewhere, review, then send cues to Sume to render. DeepL's API terms and pricing are between you and DeepL.
That boundary is worth stating plainly because several products bundle translation and subtitles under one price. Here you pay Sume only for the transcription and the render, and you pick the translation engine independently, which also lets you switch it per language pair.
How should translated cues be written?
Translated text rarely has the same length as the source, so a card that fit in English can wrap or overflow in another language. Sume does not translate, so it does not resize for you. Use the design.phrasing fields (max_words, max_chars, pause_seconds) to control how a card breaks, and cut cues so each stays readable. Our post on 42 characters per line shows how to set the limit.
Check every language on a short clip before the batch: render the longest translated line you have, watch where the card breaks and adjust max_chars for that language only. Keep the look otherwise identical across languages so the set feels like one campaign.
Pick the style by wording, not by language code. The caption language field is only a speech-to-text hint. Korean copy needs a Hangul style such as black-outline; sending it to slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style.
| Question | DeepL Voice | Sume captions |
|---|---|---|
| When | During the call | After, on a recording |
| Who translates | DeepL | You, or a tool you choose |
| Output | On-screen subtitles in the meeting app | A new MP4 with burned-in text |
| Languages | 40+ listed | Wording you supply; Latin and Hangul styles documented |
| Price shape | See DeepL pricing | $0.20 per job up to 60 seconds |
A recommended split
If your team runs customer calls across languages, DeepL Voice or a similar live tool is the right layer for conversations. If you then publish a clip of the call (a testimonial, a training moment), take the recording, translate the transcript with the tool you trust, review it, and burn it with Sume. Run one job per language so each output has its own URL, and send a stable Idempotency-Key per language so retries cannot double-charge. The bilingual subtitles post covers stacking two languages on one clip.
Sources
Related posts
More in Comparisons
- Descript embedded SRT toggle vs Sume burned-in captions
Descript added a toggle for the embedded SRT track on export. Sume's caption job burns text into pixels and does not accept SRT uploads. When each fits.
- Descript music at a set length vs Sume music length in the prompt
Descript's Sept 17 update generates music and effects at specified lengths. Sume Music has no duration field; you ask for length in the prompt. How to do that.
- Descript per-second smoothing credits vs flat API job prices
Descript now bills smoothing by the second. Here is how that compares with Sume media jobs, which charge a flat amount per job or per output minute.
- Dreamina Long Video Mode: 3 minutes vs Sume's 30 s Seedance
Dreamina says its Seedance 2.5 Long Video Mode makes videos up to three minutes. Sume's seedance-2.5 makes 4-30 s per request, so long films are joined clips.
Written by Sume