Dialogue in Korean, Japanese or Spanish: Kling 3.0 vs Seedance 2.5
Which spoken languages Kling 3.0 Omni and Seedance 2.5 list, what Sume's generate_audio flag does, and how to test a Korean line before a full render.

Kling 3.0 Omni lists five dialogue languages (English, Chinese, Japanese, Korean and Spanish), and the Dreamina page for Seedance 2.5 lists eleven (Chinese, English, Spanish, Indonesian, Malay, Thai, Arabic, Portuguese, Vietnamese, Japanese and Korean). Both vendors describe speech generated with the picture. Sume has a generate_audio flag but no language field, so you choose the language by writing the line in that language.
One caution before the details: the Kling list is for Kling 3.0 Omni and the Seedance list is from Dreamina's app page. Sume's catalog id is kling-3 and seedance-2.5, and Sume does not publish its own language list. Treat the vendor lists as the best available hint and test your language.
Which languages does each vendor say it supports?
Kling's Omni guide says native audio is generated in the same pass as the video, with no separate lip-sync step, and that it supports English, Chinese, Japanese, Korean and Spanish, with American, British and Indian English accents and several Chinese dialects. It also says users can specify what each character says, how, and when.
The Dreamina Seedance 2.5 page lists multilingual support in eleven languages and says lip-sync is available for audio generations. ByteDance's own Seedance 2.0 launch post describes lip-synced multilingual dialogue and sound effects generated jointly with the video, without naming languages.
| Language | Kling 3.0 Omni guide | Dreamina Seedance 2.5 page |
|---|---|---|
| English | Yes, with US, UK and Indian accents | Yes |
| Chinese | Yes, with several dialects | Yes |
| Japanese | Yes | Yes |
| Korean | Yes | Yes |
| Spanish | Yes | Yes |
| Portuguese, Thai, Vietnamese, Indonesian, Malay, Arabic | Not listed | Yes |
What does Sume let you control?
The video request has a generate_audio boolean that defaults to the model's audio capability, and the video docs list Seedance 2.x as honoring audio references. There is no language, voice or accent field on the generation request. The language comes from the words in your prompt, and any accent hint is just more prompt text that the model may or may not follow.
Whether Sume's kling-3 row is the same model as the Omni guide's feature set is not something the docs claim. Sume's Kling row takes text or start/end frames and rejects reference inputs, so Omni features such as element references are not available through it. Check generate_audio in GET /v1/videos/models before relying on sound.
How do you test a Korean, Japanese or Spanish line cheaply?
Write the spoken line in the target language, inside quotation marks, and say who speaks it. Keep it to one short sentence for the first test, because a 4 second clip is enough to hear whether the pronunciation and mouth shapes are plausible. Run it at 480p on a Seedance row (the 480p tier is the cheapest) and listen before you spend on 1080p.
Keep the camera still while you judge lip-sync: a locked medium close-up gives the model the simplest job. If the clip passes, rerun the same prompt at the final resolution and length.
{
"model": "seedance-2.5",
"prompt": "Medium close-up, locked camera, a woman at a cafe table looks at the lens and says in Spanish: \"Hola, el café de hoy es suave y dulce.\" Quiet cafe ambience.",
"duration": 5,
"resolution": "480p",
"aspect_ratio": "9:16",
"generate_audio": true
}What if the audio is wrong or you need exact subtitles?
Generated speech can drift from the words you wrote, so do not treat the audio track as a script of record. For subtitles that must match your copy exactly, burn them in afterwards: the video captions docs take a finished video_url and either transcribe the speech or accept authored cues with text, start and end, which skips speech-to-text. A language value there is only a speech-to-text hint, and it never selects the caption style.
If the speech is wrong in a way no prompt fixes, regenerate with generate_audio: false and add narration separately, or pick another row. Sume does not guarantee a particular accent or voice, and a failed take is still a captured job, so keep drafts small.
Sources
Related posts
More in Models
- Which AI video model for a 15-second single take on Sume?
Kling 3.0, Wan 3.0, MiniMax H3 and Seedance 2 reach 15 seconds on Sume; Auto and Grok stop at 10. Seedance 2.5 and Wan 3.0 go to 30. Table of limits.
- ByteDance MPA copyright pact: what Seedance users should check
ByteDance signed a copyright agreement with the MPA on Aug 17, 2026 covering Seedance. What was disclosed, what was not, and how to handle a refused request.
- Can GPT-6 Sol, Luna or Astra generate images or video?
No. OpenAI lists GPT-6 Astra, Sol and Luna as text-output models with image input. Where image and video generation live instead, and how Sume splits it.
- Claude Haiku 4.5 retires no sooner than Oct 15: Sume's Haiku row
Anthropic lists Haiku 4.5 retirement as not sooner than October 15, 2026. What Sume's Haiku 4.5 row does, what it has no successor for, and how to move off.
Written by Sume