Thai, Vietnamese, Indonesian voiceover: on MAI's list, not Sume's tags
MAI-Voice-2.1 lists th-TH, vi-VN and id-ID; Sume's 16 voice tags do not. What that means, a 1-cent audition, and the Unicode trap in Vietnamese.

Microsoft lists Thai, Vietnamese and Indonesian for MAI-Voice-2.1 and its Flash sibling. Sume's voice library does not tag any of the three, so on Sume treat them as untested: send one cent of audition before you plan a campaign around them.
What Microsoft lists
Microsoft's model page includes Thai, Vietnamese and Indonesian among its 23 languages, and the OpenRouter listing for the Flash model carries th-TH, vi-VN and id-ID codes. These are the clearest cases where the launch adds coverage that Sume's own voice list does not claim.
| Locale code | Listed on | Note |
|---|---|---|
| th-TH | MAI-Voice-2.1 and Flash | On Microsoft's model page and the OpenRouter locale list |
| vi-VN | MAI-Voice-2.1 and Flash | Same |
| id-ID | MAI-Voice-2.1 and Flash | Same |
| th, vi, id | Sume voice tags | None of the three is among Sume's 16 tags |
What Sume ships
Sume's library tags 16 languages and these three are not among them. The language field still accepts any 2 to 16 character string, and when a voice's language metadata is unknown the request is not blocked by the mismatch guard. That is not a promise that every voice reads Thai well. This post makes no quality claim; it tells you how to find out for a cent.
Sume's TTS Router (POST /v1/tts-router/generate) takes a required model (sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview), one of transcript or transcript_source, a voice selector and a language string of 2 to 16 characters. It is character-metered at $0.0475 per 1,000 characters, rounded up to whole cents per job, with a 20,000-character cap per request.
A one-cent audition
Write one line of about 200 characters that contains your hardest words: a price, a date, a brand name and a number. Save it as line.txt, then submit it and read the job. A line of 210 characters or fewer is the minimum one-cent job. Set language to the code you will use in production; if the chosen voice is tagged for a different language you get the 409 guard rather than a charge.
JOB=$(curl -sS -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audition-th-001" \
-d "$(jq -n --rawfile t line.txt '{model:"sonic-3.6",transcript:$t,avatar_handle:"product_host",language:"th"}')" \
| jq -r '.data.request_id')
curl -sS https://api.sume.com/v1/jobs/$JOB/status -H "Authorization: Bearer $SUME_API_KEY"
curl -sS https://api.sume.com/v1/jobs/$JOB/result -H "Authorization: Bearer $SUME_API_KEY"What to listen for
- Thai has no spaces between words, so check that phrase breaks in your script fall where a listener would pause.
- Vietnamese tone marks must survive your pipeline. Normalise text to NFC before you send it (see the code below).
- Indonesian numbers and abbreviations: check how a price, a percentage and a short form such as
ygare read. - Brand names in Latin script inside Thai or Vietnamese text.
- If nothing sounds right, stop at the audition. You spent one cent and nothing else.
Normalise Vietnamese before you count or send
A composed letter such as the one in the word Viet is one Unicode code point in NFC but three in NFD. The same sentence can differ by several characters between the two forms, which changes both what you are billed and what the engine receives. The snippet prints the three lengths for one phrase.
import unicodedata
s = "Ti\u1ebfng Vi\u1ec7t"
print(len(s))
print(len(unicodedata.normalize("NFC", s)))
print(len(unicodedata.normalize("NFD", s)))What it costs
A 300-character line is a 2-cent job; your audition lines are 1 cent each when they are 210 characters or fewer.
| Test | Characters per line | Jobs | Sume cost |
|---|---|---|---|
| One line | 210 or fewer | 1 | $0.01 |
| 6 lines (one per hard case) | 210 or fewer | 6 | $0.06 |
| One 1,000-character script | 1,000 | 1 | $0.05 |
Where this stops
Sume's TTS is asynchronous: you submit, poll GET /v1/jobs/:id/status, then read /result. It does not stream audio and it has no SSML field, so pauses and emphasis come from your punctuation and the generation_config controls (speed 0.6 to 1.5, volume 0.5 to 2, a short emotion guide). Microsoft's Flash model is the one built for live conversation; for a rendered ad or a narrated video, a finished file is the thing you need.
Sources
Related posts
More in Models
- Translate text in an image but keep brand names: the edit prompt
A prompt pattern for translating the words in an image on Sume while leaving brand names, prices and codes alone, with an Ideogram 4.5 request.
- Turkish TTS API: MAI-Voice-2.1 tr-TR vs Sume tr, and the dotted I bug
Turkish is on MAI-Voice-2.1 (tr-TR) and in Sume's voice tags (tr). Upper-casing a script in code turns i into I; the failure and a one-cent test to catch it.
- /v1/videos audio_url references: Seedance 2 yes, Omni and Kling no
Only models whose supported_input_references lists audio_url take audio references on Sume: Seedance 2.x, Wan 3.0 and MiniMax H3 yes; Omni and Kling no.
- Veo 3.1 1080p needs 8 seconds; Omni on Sume does 3 to 10 at 1080p
Veo 3.1 only renders 1080p and 4K at 8 seconds. Sume's Omni row takes any length from 3 to 10 seconds at 1080p, so a 5-second clip costs $0.94 billed.
Written by Sume