Kling 4.0 lip-sync reaches nine languages: a one-clip test each
Kling 4.0 adds Portuguese, German, French and Hindi lip-sync to five existing languages. A short per-language test plan to run before you promise localization.

Kling 4.0 lists nine lip-sync languages: Portuguese, German, French, and Hindi were added to Chinese, English, Japanese, Korean, and Spanish, according to the Kling 4.0 vs 3.0 page. A language on the list tells you it is supported, not that it looks right. Test one short clip per language before you promise a localized campaign.
The nine languages
The same page lists improved lip-sync and two-channel stereo audio for 4.0. Nothing on it quantifies quality, so the test below is the only way to learn how a language performs for your voice, script, and framing.
A test matrix
Use the same shot, the same speaker description, and one sentence of comparable length per language. Differences you see are then about the language, not the setup.
| Language | Status on Kling's list | Check by eye | Check by ear |
|---|---|---|---|
| Chinese, English, Japanese, Korean, Spanish | Existing languages | Mouth shapes on stressed syllables | Timing against the mouth |
| Portuguese, German, French, Hindi | Added in 4.0 | Mouth shapes on stressed syllables | Timing against the mouth |
Keep tests short and cheap
Keep each test to 5 to 8 seconds so a failed take is cheap. Sume reserves the estimated cost when a request is accepted and releases the reservation for failed jobs, per the admission docs, but a completed bad take is still billed. The docs name no Kling 4.0 id yet, so list the live catalog and check generate_audio and supported_durations before you plan around it.
What to score
Scoring keeps a language test from becoming opinion. Use three questions per clip and answer each yes or no: does the mouth open for every stressed syllable, does the speech start and stop with the mouth, and are names and numbers pronounced correctly. A language passes when the answer to all three is yes on both takes.
Add one native speaker to the review if you can. A non-speaker can judge mouth timing, but only a speaker can judge whether the pronunciation sounds natural, and that is often what a local audience notices first.
- Mouth opens on stressed syllables.
- Speech starts and stops with the mouth.
- Names and numbers are correct.
- A native speaker confirms naturalness.
Record what failed
Write down, per language, what failed: wrong mouth shapes, drift across the clip, or mispronounced names. Brand names and numbers are the usual trouble, so include one in every test sentence. The video docs describe the request fields. If a language fails, plan a fallback route such as a voiceover track laid under a silent clip.
- Use one sentence with a brand name and a number in every language.
- Run each language twice to see take-to-take variance.
- Keep the best clip and the failure notes with the job id.
- Decide a fallback before launch, not after a client review.
Sources
Related posts
More in Use cases
- LinkedIn video ad ratios: 16:9, 1:1, 4:5, 9:16 pixel sizes, one table
LinkedIn video ads accept 16:9, 1:1, 4:5 and 9:16 with different pixel rules and under 30 fps. See the table and how to match each ratio on Sume.
- Meta ad video length by placement: 15 s to 240 min
Meta's placement chart sets different lengths and ratios per placement, from 15 s Messenger Stories to 240 min Feed. A render plan by placement.
- Meta auto-detects AI-made ads: plan for the AI info label
Meta says its ad system detects ads made or edited with third-party AI tools through industry-standard signals and may show an AI info label. What to prepare.
- Meta labels AI images from C2PA and IPTC metadata: what stripping does
Meta reads C2PA and IPTC invisible metadata to label AI images. What the February 2024 post says, what it leaves open for video, and how to check files.
Written by Sume