Kling 3.0 dialogue languages: is French or German covered?
Kling lists Chinese, English, Japanese, Korean and Spanish for VIDEO 3.0 dialogue. French and German are not on that list, so test a 720p clip on Sume first.

Which languages does Kling 3.0 list for dialogue?
Five, on Kling's VIDEO 3.0 versus Omni 3.0 page: Chinese, English, Japanese, Korean and Spanish, plus dialects and accents. French and German are not named there. Kling's Kling 4.0 versus 3.0 page puts the language count at 5 for Kling 3.0 and "10+" for Kling 4.0.
If your script is in French or German, the vendor pages do not promise it for 3.0. That does not prove it fails; it means you should not assume it works.
| Source page | Model | Language statement |
|---|---|---|
| VIDEO 3.0 vs Omni 3.0 | VIDEO 3.0 | Chinese, English, Japanese, Korean, Spanish, plus dialects and accents |
| Kling 4.0 vs 3.0 | Kling 3.0 | 5 languages |
| Kling 4.0 vs 3.0 | Kling 4.0 | 10+ languages |
Does Sume let me choose a dialogue language?
No. The /v1/videos request has no language field. Audio on kling-3 is the optional generate_audio boolean, and the language of speech comes from the words in your prompt. There is also no voice sample input on this row.
Sume does not publish a language list for the model, so the vendor list above is the only guidance and Sume adds no guarantee to it.
How do I test a language cheaply?
Run a short clip at the lowest resolution the row offers. kling-3 starts at 720p and 4 seconds, so a 4 second, 720p clip with audio on is the cheapest meaningful test. Put one short line in quotes in the prompt, then listen.
Each submit reserves the provider list price times 1.25 from the workspace balance, and the finished job reports usage.cost, so you know the test cost afterwards and can compare audio on versus off in the audio cost comparison.
{
"model": "kling-3",
"prompt": "A baker looks at the camera and says in French: \"Bonjour, le pain est chaud.\" Soft morning light, static camera.",
"duration": 4,
"resolution": "720p",
"aspect_ratio": "9:16",
"generate_audio": true
}How do I check what was actually said?
Do not trust your ear alone on a language you do not speak. The guide on checking dialogue by transcript shows how to inspect the finished video: read the transcript back and compare it to the line you wrote, including the words and the language.
If the line comes back in the wrong language or garbled, rewrite the prompt with the language named and the line shorter, and run it again.
What if Kling 3 fails the test?
Try a model that lists your language, or render silent and add speech separately. Seedance rows take generate_audio too and seedance-2.5 also accepts an audio reference, which is a different route to a voice. Kling 4.0 is described on Kling's pages as covering more languages, but it is not a Sume model id today. Pin kling-3 for repeatable tests and re-check the catalog when a new Kling id appears.
Sources
Related posts
More in Models
- Kling 3.0 shot timing prompts: make seconds add up on Sume
Kling's guide writes multi-shot prompts as 'Shot 1, 4s'. On Sume kling-3 has no shot list field, so the timing lives in the prompt and in duration.
- Kling 3.0 subject binding limits and what Sume kling-3 takes
Kling's subject binding takes up to 4 images, a 3-8 s clip and a voice sample. Sume's kling-3 takes none of them, so lock identity with a first frame.
- Kling 3.0 in the Sume Videos panel: 5-15 s choices vs API 4-15 s
The panel offers Kling 3.0 at 5, 6, 8, 10 or 15 seconds and no end frame. The API accepts any whole second from 4 to 15 and a last frame.
- Kyutai Pocket TTS languages: what it speaks vs hosted Sume TTS
Pocket TTS is a 100M-parameter open model you run yourself. Its README lists seven languages; Sume's TTS is hosted, per character, with a voice id.
Written by Sume