MiniMax H3 languages: which spoken languages does it support?

MiniMax says H3 has stable support for 11 dialogue languages and partial support for others. Sume's docs list no language set, so test yours first.

4 min readSume
All posts

MiniMax H3 has stable support for 11 dialogue languages, according to MiniMax: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian and Spanish. Other languages work "to varying degrees". Sume's docs do not list a language set for minimax-h3, so treat MiniMax's list as a guide and test your language.

As of 2026-09-29, the language list is on MiniMax's open-source announcement, in a specification table for the model. Sume's side comes from Video generation, which describes minimax-h3-max audio as native stereo and names no languages.

Which languages does MiniMax list?

From MiniMax's H3 open-source announcement, read 2026-09-29.
ItemWhat MiniMax's page says
Stable dialogue languages11: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish
Other languages"Additional languages are also supported to varying degrees"
Output audio32 kHz stereo

Does Sume say which languages H3 speaks?

No. The Sume docs describe minimax-h3-max as returning native stereo audio and say both H3 ids accept audio references, but they do not name languages, and this post does not add any. The MiniMax page describes its model release; it is not a promise about the hosted route Sume calls.

How do I test my language cheaply?

Run one short clip at the lowest resolution and listen. On Sume, minimax-h3 takes 5 to 15 seconds at native 480p or 768p, so a 5-second clip at 480p is priced at about $0.3125 (list × 1.25, plus a 5.5% agent fee by default). Write the line you want spoken into the prompt. Repeat the test once for each language you plan to ship, and keep the spoken line short so the result is easy to judge.

curl -X POST "https://api.sume.com/v1/videos" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: h3-language-test-001" \
  -d '{
    "model": "minimax-h3",
    "prompt": "A woman at a cafe table says a short greeting to the camera, natural room sound",
    "resolution": "480p",
    "duration": 5
  }'

Can I match a specific voice instead?

Sume's docs say audio references are honored by minimax-h3 and minimax-h3-max. A reference clip is the documented way to steer sound; see MiniMax H3 voice reference. It does not change which languages the model supports.

Do Sume's limits match MiniMax's page?

Not on clip length. MiniMax's specification table gives 4–15 seconds for the open release, while Sume's docs list 5–15 seconds for minimax-h3. Send 5 seconds or more, and trim afterwards if you need a shorter clip. Read each limit on the page of the product you call.

What if my language is not on the list?

Test it the same way, and judge the result yourself. Run the short clip above with the line written in your language, listen for the words you asked for, and only then plan a batch. MiniMax says other languages are supported to varying degrees, without naming them. For a fixed script, generating the speech separately is another route; AI voiceover for short videos covers it.

Sources

Related posts

More in Models

All Models posts

Written by Sume