TikTok Translate and Dub: lip-sync limits and the AI label

TikTok's Translate and Dub in Symphony Creative Studio can clone the voice and lip-sync, but not for crowds or covered mouths. Texas and Illinois are excluded.

5 min readSume
All posts

Translate and dub in TikTok's Symphony Creative Studio changes your video's voiceover into new languages, lets you clone the original voice or pick a stock voice, and offers lip-sync, which TikTok says does not work for videos with several speakers, covered lips or mouths, poor lighting, or background noise. It is not available in Texas and Illinois, and every exported video gets an AI-generated label. Sume does not dub existing footage in anything its docs describe, so this post is about reading TikTok's limits correctly before you plan a campaign around them.

TikTok's side comes from How to Translate and dub videos with Symphony Creative Studio, read on 2026-10-02. Sume's side comes from Models and Video captions.

What does Translate and dub offer?

TikTok's page describes three choices. You can clone the voice of the original video or select a stock voice. You can turn on lip-sync so the mouth matches the dubbed audio. And you can replace the original subtitles with translated ones.

The page does not list the supported languages, video length limits, or file size limits, so those are unknown until you open the tool.

Symphony Translate and dub, from TikTok's help page, read 2026-10-02.
TopicWhat the page says
VoiceClone the original voice or choose a stock voice
Lip-syncAvailable, but not for several speakers, covered lips or mouth, poor lighting, or background noise
SubtitlesOriginal subtitles can be replaced with translated ones
Edits to the scriptMinor corrections only, such as typos; changes that alter the substance, tone, or message are not permitted
LabelAn AI-generated label is added to all exported videos
Excluded regionsTexas and Illinois in the United States
ReviewContent goes through moderation before it is published on TikTok or in Ads Manager

Why do the lip-sync limits matter for ads?

Many ad clips break those conditions on purpose: a crowd scene, a hand over the mouth, a busy cafe, or a dim review shot. If your source clip has any of them, plan for the dub to run without lip-sync, which means the audio will not match the mouth shapes.

The edit rule also matters. TikTok allows minor typo fixes to the generated script, but not changes that alter the message. A localized ad that needs a different offer per market cannot be reworded inside the tool; build that variant separately.

What does the AI label mean for your own captions?

Because the label is added on export, you do not add your own for the dubbed audio inside that tool. If you assemble a dub elsewhere, the label question is yours. See TikTok AIGC label or your own caption for how a burned notice fits.

What can Sume do instead?

The models page lists VEED Fabric 1.0 and MiniMax H3 Max Lip Sync, and both take a still image plus an audio file. The same page says video models do not lip-sync to generated speech or to a later voice-over. So the docs describe talking-face generation from a still, not re-syncing the mouth in a clip you already shot.

For a real-person clip, a Sume-side localization means new audio plus burned captions: video captions burns translated cues you supply, at $0.20 per job for a video up to 60 seconds, without touching the voice. The result is a translated caption track over the original speech, not a dub. If you want the original voice replaced and the lips matched, Symphony's tool is the one that describes doing that, within its stated limits.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume