Vietnamese, Thai, Indonesian, Malay TTS API: vi, th, id, ms on Sume

Cartesia Sonic 3.6 lists vi, th, id and ms. How to send each as the language on Sume TTS, why the field is required, and a per-character cost for each script.

4 min readSume
All posts

Southeast Asian narration works on Sume TTS with one change from an English request: set language to vi (Vietnamese), th (Thai), id (Indonesian) or ms (Malay) on POST /v1/tts-1.0/generate. Cartesia lists all four on Sonic 3.6, Sonic 3.5 and Sonic 3, and Sume forwards the field to the engine unchanged. Without it the provider default is English, and Sume's own script inference only covers Korean and Japanese.

Vietnamese and Indonesian are written in Latin letters, which makes the missing-field mistake easy to miss: the request succeeds and the audio is wrong. Thai is the one with a different script and no spaces between words.

Which codes and models cover these languages?

None of the four is in Sume's voice-library language list, so you cannot create and save a voice in these languages through Sume's voice tool. You can still synthesize with a voice UUID you hold. Sume treats Assets and Voices as optional references, and the library never gates a raw voice id.

Language codes per Cartesia's model pages (read 2026-10-03); the last column is Sume's voice-library list from its voice code
LanguageCodesonic-3.6-2026-08-27sonic-3.5-2026-05-04sonic-3-2026-01-12In Sume's voice-library language list
VietnameseviYesYesYesNo
ThaithYesYesYesNo
IndonesianidYesYesYesNo
MalaymsYesYesYesNo

Why is language required even for Latin-script text?

The OpenAPI contract describes language as the language the voice speaks the transcript in, and tells callers to set it for every non-English transcript: omitted means English at the provider. Indonesian text sent as English is pronounced with English letter rules. Sume's fallback inference only recognizes a mostly-Hangul transcript as ko and a mostly-kana one as ja; it does nothing for Vietnamese, Thai, Indonesian or Malay.

The same field is also what Sume compares with a saved voice's language. A mismatch returns HTTP 409 tts_voice_language_mismatch before a job exists, so a voice saved as en will tell you if you aim it at id.

What does a Thai or Vietnamese script cost?

Sume bills per transcript character at $0.0475 per 1,000 characters, spaces and punctuation included. Thai uses few spaces, so a Thai paragraph packs more characters than the same English paragraph has words; count characters, not words, when you budget. The limit is 20,000 characters per request and 1,200 seconds of audio per job, whichever you hit first.

Cost at $0.0475 per 1,000 characters, from Sume's published TTS price (Sume API OpenAPI document, read 2026-10-03)
Transcript length (characters)Cost
1,000$0.0475
5,000$0.2375
20,000 (request maximum)$0.95

How do I send the request?

Set VOICE_ID to a voice that speaks the target language and change the two lines that matter, transcript and language.

import os
import uuid

import requests

r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={
        "x-api-key": os.environ["SUME_API_KEY"],
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "transcript": "Xin chào, đây là bản tin hằng tuần của chúng tôi.",
        "language": "vi",
        "voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
        "mode": "async",
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

How do I check the first take?

Queue a short sample, around 200 characters, before the full script. At $0.0475 per 1,000 characters a 200-character sample costs less than one cent ($0.0095), and it tells you whether the voice reads tones, diacritics and loanwords the way you want. If the pronunciation of a brand name is off, spell it phonetically in the transcript or use the pronunciation dictionary id field. Only then submit the long script, and keep the same language code for every part so the takes match.

Sources

Related posts

More in Models

All Models posts

Written by Sume