MAI-Voice-2.1 lists Korean, Thai, Vietnamese: Sume's language field
Microsoft's MAI-Voice-2.1 page lists 23 languages including Korean, Thai and Vietnamese. What to check before you plan a non-English voiceover on Sume TTS.

Yes: Microsoft's MAI-Voice-2.1 page lists Korean, Thai, Vietnamese and Turkish among its 23 languages (read 2026-10-04). Sume's TTS does not publish a language table in the docs read for this post, so the safe plan for a non-English voiceover on Sume is to set the language field on every request, send a short test line first, and listen before you commit a whole script. Below is the list Microsoft gives and what Sume's request expects.
What Microsoft lists
Microsoft's page states 23 languages and lists regional variants of English, Spanish and Portuguese. The model page's list, grouped for reading, is below.
| Group | Languages listed |
|---|---|
| English variants | US, Australia, UK, India |
| Western Europe | Italian, French, German, Spanish (Spain, Mexico), Portuguese (Brazil, Portugal), Dutch |
| Northern and Central Europe | Swedish, Norwegian (Bokmal), Danish, Finnish, Polish, Czech, Hungarian, Romanian |
| Eastern Europe and Turkey | Russian, Turkish |
| Asia | Hindi, Korean, Chinese (Simplified), Thai, Vietnamese, Indonesian |
What Sume's request needs
The API reference and the OpenAPI file describe POST /v1/tts-1.0/generate with a language field: a BCP-47 or ISO-639 code such as ko, ja or en. The schema says to set it for every non-English transcript, because an omitted value defaults to English at the provider. Sume infers Korean or Japanese from a transcript written only in Hangul or kana as a fallback, but do not rely on that for mixed text.
Two practical consequences follow. First, a Thai or Vietnamese script sent without language is a request you should expect to sound wrong, not a request the API will rescue. Second, the voice also matters: a voice selected through avatar_id or avatar_handle was built for a language, and a mismatch triggers a warning you confirm with confirm_language_mismatch. The failure modes are covered in the mismatch warning post and the Korean script error post.
A test plan before you commit
Do this for each target language, not once.
- Pick one 150-character line with a number, a brand name and a question mark, translated by a native speaker.
- Send it with
languageset and the voice you intend to ship. - Listen for stress on the brand name and for how the number is read aloud.
- Only then queue the full script, which can be up to 20,000 characters per job.
Where this leaves the choice
If your language is on Microsoft's list, MAI-Voice-2.1 is priced at $22 per million characters on its page (read 2026-10-04). If you need the speech inside a pipeline with job URLs, hosted artifacts, captions and a timeline, Sume's job model is the point; see the models overview for how those pieces connect. OpenAI's text-to-speech guide says its voices are currently optimized for English while listing a long set of supported languages, which is a different promise from a per-language list.
For the wider picture of voice identity across languages, read MAI-Voice-2.1: 23 languages and voice matching versus Sume's voice id.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 or Flash for explainer narration: latency barely matters
Microsoft lists about 550 ms vs 45 ms model inference. For narration rendered ahead of time, pick on quality and price, then see how Sume jobs return audio.
- Meshy 7.1 Ultra 4K: when a 3D model is overkill
Meshy 7.1 added Ultra 4K meshes of up to 80 million triangles. Sume has no image-to-3D model; to show a product turning, a short video does it.
- Midjourney has no public API: comparable control in a pipeline
Midjourney's 10/1 alpha adds a pinnable --exp setting and edits that keep aspect ratio. What to use when you need that kind of control from code.
- Midjourney moodboards now take uploads: reference sets for Sume edits
Midjourney's October 1 changelog lets moodboard pages accept file uploads. To reuse a reference set in an API, store the URLs and send them as input_references.
Written by Sume