MAI-Voice-2.1 lists Korean, Thai, Vietnamese: Sume's language field

Microsoft's MAI-Voice-2.1 page lists 23 languages including Korean, Thai and Vietnamese. What to check before you plan a non-English voiceover on Sume TTS.

5 min readSume
All posts

Yes: Microsoft's MAI-Voice-2.1 page lists Korean, Thai, Vietnamese and Turkish among its 23 languages (read 2026-10-04). Sume's TTS does not publish a language table in the docs read for this post, so the safe plan for a non-English voiceover on Sume is to set the language field on every request, send a short test line first, and listen before you commit a whole script. Below is the list Microsoft gives and what Sume's request expects.

What Microsoft lists

Microsoft's page states 23 languages and lists regional variants of English, Spanish and Portuguese. The model page's list, grouped for reading, is below.

MAI-Voice-2.1 languages named on Microsoft's model page (read 2026-10-04)
GroupLanguages listed
English variantsUS, Australia, UK, India
Western EuropeItalian, French, German, Spanish (Spain, Mexico), Portuguese (Brazil, Portugal), Dutch
Northern and Central EuropeSwedish, Norwegian (Bokmal), Danish, Finnish, Polish, Czech, Hungarian, Romanian
Eastern Europe and TurkeyRussian, Turkish
AsiaHindi, Korean, Chinese (Simplified), Thai, Vietnamese, Indonesian

What Sume's request needs

The API reference and the OpenAPI file describe POST /v1/tts-1.0/generate with a language field: a BCP-47 or ISO-639 code such as ko, ja or en. The schema says to set it for every non-English transcript, because an omitted value defaults to English at the provider. Sume infers Korean or Japanese from a transcript written only in Hangul or kana as a fallback, but do not rely on that for mixed text.

Two practical consequences follow. First, a Thai or Vietnamese script sent without language is a request you should expect to sound wrong, not a request the API will rescue. Second, the voice also matters: a voice selected through avatar_id or avatar_handle was built for a language, and a mismatch triggers a warning you confirm with confirm_language_mismatch. The failure modes are covered in the mismatch warning post and the Korean script error post.

A test plan before you commit

Do this for each target language, not once.

  • Pick one 150-character line with a number, a brand name and a question mark, translated by a native speaker.
  • Send it with language set and the voice you intend to ship.
  • Listen for stress on the brand name and for how the number is read aloud.
  • Only then queue the full script, which can be up to 20,000 characters per job.

Where this leaves the choice

If your language is on Microsoft's list, MAI-Voice-2.1 is priced at $22 per million characters on its page (read 2026-10-04). If you need the speech inside a pipeline with job URLs, hosted artifacts, captions and a timeline, Sume's job model is the point; see the models overview for how those pieces connect. OpenAI's text-to-speech guide says its voices are currently optimized for English while listing a long set of supported languages, which is a different promise from a per-language list.

For the wider picture of voice identity across languages, read MAI-Voice-2.1: 23 languages and voice matching versus Sume's voice id.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume