Chinese, Hindi and Russian text to speech API: Sume zh, hi and ru

Sume's Voices library has zh, hi and ru tags. How to request each, what Eleven v4 lists, and the one field that stops an English-sounding read.

4 min readSume
All posts

For Chinese, Hindi or Russian text to speech on Sume, use a Voices library voice tagged zh, hi or ru and send the same code as language on the request. Those are three of the 16 voice language tags Sume lists. The Eleven v4 docs page names Mandarin Chinese, Hindi and Russian too (read 2026-10-03), so both vendors cover them on paper.

What the tags mean here

Sume's tag is a two-letter code, so Chinese is zh with no regional split in the tag. Sume's 16 tags have no regional variants, so test a sample if you need a particular dialect.

Language names as listed by each vendor, read 2026-10-03
You needSume tagEleven v4 docs list
Mandarin ChinesezhMandarin Chinese
HindihiHindi
RussianruRussian

The request

The tts_create body takes the transcript (up to 20,000 characters), a voice.id, and language. Without language the provider assumes English according to the schema, and Sume infers only Korean and Japanese from the script of the text. Hindi, Chinese and Russian get no such fallback, so always set the field. Spaces and punctuation count toward the $0.0475 per 1,000 characters rate, as listed in the Sume catalog.

Price is by characters you send, so count your real transcript instead of estimating from word counts.

Captions that match

The caption guide documents a rule only for Korean: Hangul copy on a Latin style such as slam is rejected. It says nothing about Chinese, Hindi or Russian glyphs, so burn one test clip before a batch; see the video captions guide.

Sources

Related posts

More in Models

All Models posts

Written by Sume