Hinglish text to speech API: romanized Hindi with language hi
For romanized Hindi or Hinglish on Sume TTS, send language `hi` even though the text is Latin script, and pin a Sonic model id on the Router.

For a Hinglish or romanized Hindi transcript, set language to hi on the Sume TTS request even though the text is written in Latin script. Sume's schema asks you to set language for every non-English transcript, and Cartesia's guide gives the same instruction for Sonic 3.6. Send the model id sonic-3.6 through the TTS Router.
What does Cartesia say about Latin-script Hindi?
Per the Sonic 3.6 page (read 2026-09-30), the model has expanded support for Hindi transcripts written in Latin script and better pronunciation of Indian names and places. The advanced guide says Sonic reads Hindi and other Indic languages in Latin script, follows romanized transcripts materially better in 3.6, and tells you to write romanized text the way it is naturally typed, keeping English words in standard spelling.
| Transcript | Example | `language` |
|---|---|---|
| Pure romanized Hindi | Aapka order aa gaya hai. | hi |
| Romanized Hindi plus English | Aapka order confirm ho gaya hai. | hi |
| Devanagari plus romanized Hindi | आपका order aa gaya hai. | Listed as not supported |
How do I send it on Sume?
The Sume language field is a 2 to 16 character BCP-47 or ISO-639 code. The schema says to set it for every non-English transcript and never to translate a non-English request into English. Omitted, it defaults to English at the provider. Cartesia's guide also sets a separate normalization value for dates and digits; Sume's request has no such field, so only language carries over.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hinglish-demo-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Aapka order confirm ho gaya hai.",
"avatar_handle": "acme",
"language": "hi"
}'Which Sonic id should I pin?
The Router model field is a required pass-through catalog id with five values in the schema: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. The Latin-script improvement is a Sonic 3.6 claim, so pin sonic-3.6 rather than an alias. More on the ids in Cartesia Sonic 3.6 API model ids.
What does the voice-language warning mean here?
If the chosen voice's primary language differs from the requested one, Sume returns a double-check warning with the text "Pronunciation may sound unnatural." and creates no job or charge yet. If you pick an English-primary voice and send hi, expect that warning; confirm only after a person agrees. The loop is explained in the mismatch warning post.
Sume's docs do not list which languages each router model accepts, so listen to a sample of your own copy first. A wider look at the field is in Sonic 3.6 languages and the language field.
Sources
Related posts
More in Use cases
- Holiday ad creative by phase: four videos, one bulk run
Meta splits the holidays into discovery, deals, gifting and fresh start. Queue one video per phase as a single Sume bulk run with a spend cap on each.
- India's synthetic-media label rule: does routine editing count?
India's amended IT rules exclude routine editing, formatting and compression from synthetic media, and require a prominent label on the rest. What Sume offers.
- Instagram API: JPEG only, no MPO or JPS. Stills from a video frame
Meta's publishing guide says JPEG is the only image format and MPO and JPS are not supported. Sume video-frames returns jpeg by default, or png.
- Instagram carousel limit of 10: make 10 stills from one video
Meta's guide limits carousels to 10 images, videos or a mix, and 100 API posts per 24 hours. Sume video-frames takes up to 24 times per call, so pick 10.
Written by Sume