Vietnamese, Thai, Indonesian, Malay TTS API: vi, th, id, ms on Sume
Cartesia Sonic 3.6 lists vi, th, id and ms. How to send each as the language on Sume TTS, why the field is required, and a per-character cost for each script.

Southeast Asian narration works on Sume TTS with one change from an English request: set language to vi (Vietnamese), th (Thai), id (Indonesian) or ms (Malay) on POST /v1/tts-1.0/generate. Cartesia lists all four on Sonic 3.6, Sonic 3.5 and Sonic 3, and Sume forwards the field to the engine unchanged. Without it the provider default is English, and Sume's own script inference only covers Korean and Japanese.
Vietnamese and Indonesian are written in Latin letters, which makes the missing-field mistake easy to miss: the request succeeds and the audio is wrong. Thai is the one with a different script and no spaces between words.
Which codes and models cover these languages?
None of the four is in Sume's voice-library language list, so you cannot create and save a voice in these languages through Sume's voice tool. You can still synthesize with a voice UUID you hold. Sume treats Assets and Voices as optional references, and the library never gates a raw voice id.
| Language | Code | sonic-3.6-2026-08-27 | sonic-3.5-2026-05-04 | sonic-3-2026-01-12 | In Sume's voice-library language list |
|---|---|---|---|---|---|
| Vietnamese | vi | Yes | Yes | Yes | No |
| Thai | th | Yes | Yes | Yes | No |
| Indonesian | id | Yes | Yes | Yes | No |
| Malay | ms | Yes | Yes | Yes | No |
Why is language required even for Latin-script text?
The OpenAPI contract describes language as the language the voice speaks the transcript in, and tells callers to set it for every non-English transcript: omitted means English at the provider. Indonesian text sent as English is pronounced with English letter rules. Sume's fallback inference only recognizes a mostly-Hangul transcript as ko and a mostly-kana one as ja; it does nothing for Vietnamese, Thai, Indonesian or Malay.
The same field is also what Sume compares with a saved voice's language. A mismatch returns HTTP 409 tts_voice_language_mismatch before a job exists, so a voice saved as en will tell you if you aim it at id.
What does a Thai or Vietnamese script cost?
Sume bills per transcript character at $0.0475 per 1,000 characters, spaces and punctuation included. Thai uses few spaces, so a Thai paragraph packs more characters than the same English paragraph has words; count characters, not words, when you budget. The limit is 20,000 characters per request and 1,200 seconds of audio per job, whichever you hit first.
| Transcript length (characters) | Cost |
|---|---|
| 1,000 | $0.0475 |
| 5,000 | $0.2375 |
| 20,000 (request maximum) | $0.95 |
How do I send the request?
Set VOICE_ID to a voice that speaks the target language and change the two lines that matter, transcript and language.
import os
import uuid
import requests
r = requests.post(
"https://api.sume.com/v1/tts-1.0/generate",
headers={
"x-api-key": os.environ["SUME_API_KEY"],
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"transcript": "Xin chào, đây là bản tin hằng tuần của chúng tôi.",
"language": "vi",
"voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
"mode": "async",
},
timeout=30,
)
r.raise_for_status()
print(r.json())How do I check the first take?
Queue a short sample, around 200 characters, before the full script. At $0.0475 per 1,000 characters a 200-character sample costs less than one cent ($0.0095), and it tells you whether the voice reads tones, diacritics and loanwords the way you want. If the pronunciation of a brand name is off, spell it phonetically in the transcript or use the pronunciation dictionary id field. Only then submit the long script, and keep the same language code for every part so the takes match.
Sources
Related posts
More in Models
- What happens to the audio in each Sume video-to-video tool
Recast and Avatar Face Swap keep the source audio, Kling motion control keeps it by default, Omni always produces native audio, Genjutsu takes no audio field.
- Where to try MAI-Voice-2.1 before paying: a 10-minute listening test
Microsoft lists the MAI Playground, Copilot Audio Expressions and Foundry for MAI-Voice-2.1. Run a ten-minute listening test with a fixed script and cost it.
- Which AI video models accept an input video on Sume?
Seedance, Wan 3.0, H3, H3 Max and Gemini Omni Flash take video references; Recast, Genjutsu and Omni edit need a source video. Kling and Grok take none.
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
Written by Sume