Thai, Vietnamese, Indonesian voiceover: on MAI's list, not Sume's tags

MAI-Voice-2.1 lists th-TH, vi-VN and id-ID; Sume's 16 voice tags do not. What that means, a 1-cent audition, and the Unicode trap in Vietnamese.

4 min readSume
All posts

Microsoft lists Thai, Vietnamese and Indonesian for MAI-Voice-2.1 and its Flash sibling. Sume's voice library does not tag any of the three, so on Sume treat them as untested: send one cent of audition before you plan a campaign around them.

What Microsoft lists

Microsoft's model page includes Thai, Vietnamese and Indonesian among its 23 languages, and the OpenRouter listing for the Flash model carries th-TH, vi-VN and id-ID codes. These are the clearest cases where the launch adds coverage that Sume's own voice list does not claim.

Thai, Vietnamese and Indonesian coverage (read 2026-10-05)
Locale codeListed onNote
th-THMAI-Voice-2.1 and FlashOn Microsoft's model page and the OpenRouter locale list
vi-VNMAI-Voice-2.1 and FlashSame
id-IDMAI-Voice-2.1 and FlashSame
th, vi, idSume voice tagsNone of the three is among Sume's 16 tags

What Sume ships

Sume's library tags 16 languages and these three are not among them. The language field still accepts any 2 to 16 character string, and when a voice's language metadata is unknown the request is not blocked by the mismatch guard. That is not a promise that every voice reads Thai well. This post makes no quality claim; it tells you how to find out for a cent.

Sume's TTS Router (POST /v1/tts-router/generate) takes a required model (sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview), one of transcript or transcript_source, a voice selector and a language string of 2 to 16 characters. It is character-metered at $0.0475 per 1,000 characters, rounded up to whole cents per job, with a 20,000-character cap per request.

A one-cent audition

Write one line of about 200 characters that contains your hardest words: a price, a date, a brand name and a number. Save it as line.txt, then submit it and read the job. A line of 210 characters or fewer is the minimum one-cent job. Set language to the code you will use in production; if the chosen voice is tagged for a different language you get the 409 guard rather than a charge.

JOB=$(curl -sS -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audition-th-001" \
  -d "$(jq -n --rawfile t line.txt '{model:"sonic-3.6",transcript:$t,avatar_handle:"product_host",language:"th"}')" \
  | jq -r '.data.request_id')
curl -sS https://api.sume.com/v1/jobs/$JOB/status -H "Authorization: Bearer $SUME_API_KEY"
curl -sS https://api.sume.com/v1/jobs/$JOB/result -H "Authorization: Bearer $SUME_API_KEY"

What to listen for

  • Thai has no spaces between words, so check that phrase breaks in your script fall where a listener would pause.
  • Vietnamese tone marks must survive your pipeline. Normalise text to NFC before you send it (see the code below).
  • Indonesian numbers and abbreviations: check how a price, a percentage and a short form such as yg are read.
  • Brand names in Latin script inside Thai or Vietnamese text.
  • If nothing sounds right, stop at the audition. You spent one cent and nothing else.

Normalise Vietnamese before you count or send

A composed letter such as the one in the word Viet is one Unicode code point in NFC but three in NFD. The same sentence can differ by several characters between the two forms, which changes both what you are billed and what the engine receives. The snippet prints the three lengths for one phrase.

import unicodedata
s = "Ti\u1ebfng Vi\u1ec7t"
print(len(s))
print(len(unicodedata.normalize("NFC", s)))
print(len(unicodedata.normalize("NFD", s)))

What it costs

A 300-character line is a 2-cent job; your audition lines are 1 cent each when they are 210 characters or fewer.

Audition cost on Sume at list rate (read 2026-10-05)
TestCharacters per lineJobsSume cost
One line210 or fewer1$0.01
6 lines (one per hard case)210 or fewer6$0.06
One 1,000-character script1,0001$0.05

Where this stops

Sume's TTS is asynchronous: you submit, poll GET /v1/jobs/:id/status, then read /result. It does not stream audio and it has no SSML field, so pauses and emphasis come from your punctuation and the generation_config controls (speed 0.6 to 1.5, volume 0.5 to 2, a short emotion guide). Microsoft's Flash model is the one built for live conversation; for a rendered ad or a narrated video, a finished file is the thing you need.

Sources

Related posts

More in Models

All Models posts

Written by Sume