Which Korean voices does MAI-Voice-2.1 have, and with what styles?
MAI-Voice-2.1 lists four ko-KR voices: Grant, Harper (role styles), Haena (14 emotion styles) and Junho (12). Pair a voice with the right caption style.

Microsoft's MAI voices page lists four Korean voices: ko-KR-Grant (male) and ko-KR-Harper (female), each with six role styles, ko-KR-Haena (female) with 14 emotion styles, and ko-KR-Junho (male) with 12. All styles are set through SSML with mstts:express-as, as listed on the page read on 2026-10-04.
The four voices
| Voice | Gender | Styles |
|---|---|---|
| ko-KR-Grant | Male | 6 roles: agent, audiobook, customer_call_center, educational, narrator, neutral |
| ko-KR-Harper | Female | Same 6 roles as Grant |
| ko-KR-Haena | Female | 14 emotions: angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, neutral, regretful, relieved, sad, softvoice, surprised |
| ko-KR-Junho | Male | 12 emotions: angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, neutral, relieved, sad, softvoice |
Picking by job
For narration, a call-centre prompt or an educational read, the role-style voices (Grant, Harper) name the job directly. For an emotional ad read, Haena has the widest emotion list, and Junho is the male option with emotions, though it has no regretful or surprised style.
The page is in public preview with no SLA, so test the style you choose rather than assuming every style sounds equally natural.
After the voice: Korean captions in Sume
If the audio ends up on a video, Sume's Video captions endpoint burns captions for $0.20 per video up to 60 seconds. For Hangul text the default style is black-outline, and a korean-ad style is available. Sending a Hangul script with a Latin-only style returns caption_hangul_text_latin_style.
- cues skip speech-to-text when you already have timings.
- script_text aligns your own script to the audio.
- source_caption_id restyles an earlier caption job.
Where Sume's own TTS fits
Sume's TTS endpoint takes a plain transcript, no SSML, so mstts:express-as styles are not a Sume feature. Its catalog is Sonic models. Choose it when you want one HTTP call and a flat $0.0475 per 1,000 characters; choose MAI when you need those named Korean styles.
Sources
Related posts
More in Models
- MiniMax H3 Max: the prompt-adherence variant on Sume
fal describes MiniMax H3 Max as tuned for prompt adherence. On Sume, minimax-h3-max runs 480p to 1080p for 5 to 15 s with frames and references.
- MiniMax H3 limits: 9 images, 3 videos, 3 audio, file caps
MiniMax's H3 guide caps prompts at 7,000 characters and references at 9 images, 3 videos and 3 audio files. Cheat sheet with the Sume limits beside it.
- MiniMax H3 references on Sume: 9 images, 3 videos, 3 audio, 12 total
minimax-h3 and minimax-h3-max accept 9 images, 3 videos and 3 audio files, 12 in total, and audio cannot be the only reference. Duration rules and errors.
- Sume motion routes by length: Kling, Recast and Genjutsu
Three Sume routes drive a result from a source clip, each with a length rule: Kling takes duration_seconds, Recast follows the source, Genjutsu a range.
Written by Sume