Tamil, Telugu, Kannada, Malayalam TTS API: language codes on Sume
Sonic 3.6 lists bn, ta, te, kn, ml, mr, gu and pa beyond Hindi. How to send each on Sume TTS, what the Hinglish note does not cover, and the cost per script.

Beyond Hindi, Sume TTS can be pointed at eight more Indian languages by setting language to bn (Bengali), ta (Tamil), te (Telugu), kn (Kannada), ml (Malayalam), mr (Marathi), gu (Gujarati) or pa (Punjabi). Cartesia lists each of them on Sonic 3.6, Sonic 3.5 and Sonic 3, and Sume forwards the field unchanged. Odia (or) and Urdu (ur) are on Sonic 3.6 only; they have their own post.
Only Hindi is in Sume's own voice-library list. The other eight can be synthesized with a voice UUID you hold, but you cannot create and save a voice in them through Sume's voice tool.
Which models carry which language?
Cartesia's pages give the same answer for all nine codes, so the table is mostly a coverage check. The final column is the part that differs by product: it shows whether you can save a voice in that language through Sume.
| Language | Code | sonic-3.6-2026-08-27 | sonic-3.5-2026-05-04 | sonic-3-2026-01-12 | In Sume's voice-library language list |
|---|---|---|---|---|---|
| Bengali | bn | Yes | Yes | Yes | No |
| Tamil | ta | Yes | Yes | Yes | No |
| Telugu | te | Yes | Yes | Yes | No |
| Kannada | kn | Yes | Yes | Yes | No |
| Malayalam | ml | Yes | Yes | Yes | No |
| Marathi | mr | Yes | Yes | Yes | No |
| Gujarati | gu | Yes | Yes | Yes | No |
| Punjabi | pa | Yes | Yes | Yes | No |
| Hindi | hi | Yes | Yes | Yes | Yes |
Does the Hinglish support apply to these languages?
No. Cartesia's August 2026 Sonic 3.6 notes describe expanded support for Hindi written in Latin script and a normalization field for pairing a spoken language with a different number and date convention. That is described for Hindi and Hinglish; the page makes no equivalent promise for romanized Tamil, Telugu or the rest, so send those in their native script.
Sume has no normalization field. Its TTS body takes transcript, language, output_format, pronunciation_dict_id, generation_config, timestamps and segmentation. If a number must be read a particular way, write it out in words in the transcript.
What do the request limits mean for Indic scripts?
Sume counts transcript characters, up to 20,000 per request, at $0.0475 per 1,000. Indic scripts use combining marks, so a short-looking line can carry more characters than you expect; count the string you send. Each job is also capped at 1,200 seconds of audio, and a longer script fails with tts_duration_exceeded, so split a long Tamil chapter by section and join the WAV takes with Sume's audio timeline (up to 20 parts, gapless).
- Use
output_format.container: "wav"for takes you will join. - Use
generation_config.speedbetween 0.6 and 1.5 to fit a slot. - Pass
timestamps.words: truefor caption timing.
Request example
Tamil, with a Tamil-capable voice UUID in VOICE_ID:
import os
import uuid
import requests
r = requests.post(
"https://api.sume.com/v1/tts-1.0/generate",
headers={
"x-api-key": os.environ["SUME_API_KEY"],
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"transcript": "எங்கள் வாராந்திர செய்திக்கு வரவேற்கிறோம்.",
"language": "ta",
"voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
"mode": "async",
},
timeout=30,
)
r.raise_for_status()
print(r.json())How should I test before a long script?
Queue a short sample of around 200 characters, which costs $0.0095 at Sume's rate, and listen for conjuncts, numerals and English loanwords. Indic scripts mix local words with English terms often, and the sample shows how the voice handles that mix. If it does not read a term correctly, spell it as it should sound. Keep the same voice id across the whole series so every chapter sounds alike, and set language identically on every request.
Sources
Related posts
More in Models
- Sume image models at a $0.02 list price: Grok, Qwen, Imagen Fast
Three Sume image models list at $0.02 per image: Grok Imagine, Qwen Image and Imagen 4 Fast. How they differ on edits, ratios and image count per call.
- Veda sparse attention for MiniMax H3: 6.8x attention, 3.1x clip
Veda's sparse attention keeps 10 percent of attention work for MiniMax H3. Why 6.8x on attention becomes 3.1x per clip, and what it needs to run.
- Gemini 3.8 Flash is stable: keep model ids in config
When a vendor ships a new Flash model, a hard-coded id ages. Read Sume ids from GET /v1/catalog and treat 404 model_not_found as a signal, not a retry.
- How long can a Veo 3.1 video get? 148 seconds at 720p
Veo 3.1 extends clips 7 seconds at a time, up to 148 seconds at 720p. Sume has no extend task, so here are the longer single-clip models and how to join clips.
Written by Sume