Sync dubbing in 92 languages vs building a dub on Sume
Sync's built-in dubbing now lists 92 languages with speaker detection. Sume has no dubbing endpoint: chain transcription with a language hint, translation, TTS.

Sync's August 31, 2026 changelog says its built-in dubbing now supports 92 languages, detects speakers automatically and returns lossless audio for the final dub. Sume has no single dubbing endpoint: you assemble transcription with an optional language hint, your own translation step, and text-to-speech.
Sync facts are from its changelog; Sume facts from the OpenAPI document and the Models and Avatar videos docs, read 2026-10-01.
What changed in Sync's dubbing?
The changelog entry says dubs can be created across the App, API, Adobe Premiere Pro and DaVinci Resolve. An earlier entry describes translating a video and lip-syncing it in one step, with translation, voice cloning and lipsync handled by Sync, listed there at 29 languages. The 92 figure is the newer one.
What do I assemble on Sume?
| Step | Sume control from the docs |
|---|---|
| Transcribe | language_code: optional BCP-47 or provider hint such as en or ko; omit for auto-detect |
| Translate | Your own step; Sume docs list no translation endpoint in this flow |
| Speak | TTS with pronunciation_dict_id and generation_config for speed, volume, emotion |
| Picture | A talking-face shot is Fabric with a still plus TTS audio |
What does Sume not do here?
The Models docs say video models do not lip-sync to generated TTS or to a later voice-over. So a dub laid over an existing clip will not have re-synced lips; if you need a face that speaks the new language, generate the speaking shot from a still plus the new audio. Speaker detection is not described in the Sume pages I read, so split speakers yourself before transcribing. The speech transcription rate card lists $0.01 per audio minute.
What about Korean captions?
Avatar video captions have a rule: a Korean script with a Latin-only style such as slam, punch or tiktok-green is rejected with 400 caption_hangul_text_latin_style, so pick a Hangul style for Korean speech. For the full chain see build an AI dubbing pipeline.
Sources
Related posts
More in Use cases
- Synthesia dub speaker attribution vs Sume speech-to-text and audio
Synthesia lets Enterprise users reassign and rename speakers in a dub. Sume gives you audio detach, speech-to-text and audio joins; speaker mapping is yours.
- Synthesia Assistant styles vs Sume's scene prompt and previews
Synthesia Assistant asks for a Cinematic or Presentation delivery style before scenes generate. Sume uses scene direction plus first-frame previews instead.
- Synthesia Avatar Builder credits: 14 per option, and Sume jobs
Synthesia Avatar Builder charges 14 credits per generated option. Sume creates an avatar with one job per request via POST /v1/avatar-1.0/generate.
- Synthesia brand kit fonts vs Sume caption fonts: Hangul only
Synthesia Motion Graphics now use brand kit fonts. Sume's caption font field takes a Hangul face only, so Latin brand fonts cannot be set there.
Written by Sume