Chinese text to speech API: set language zh or it reads as English

Sume TTS only guesses Korean and Japanese when the language is missing. For Mandarin send language zh and pick a voice tagged zh, then test one line.

5 min readSume
All posts

To make Mandarin speech with Sume's text to speech API, send the Chinese text as transcript, set language to zh, and choose a voice that is tagged for Chinese. Do not leave language out: Sume's fallback only recognises Hangul and kana, so a Chinese script with no language set is handed to an English-defaulted engine.

This follows the API reference, read 2026-10-03. Sume's reference does not publish a list of TTS languages, so what follows separates what is documented from what you should verify with one test line.

What does the fallback actually guess?

The language field is documented as the language the voice speaks the transcript in, with ko, ja and en as examples. Omitted, it defaults to English at the provider. Sume adds one rescue: if the text is mostly Hangul it infers ko, and if it is mostly kana it infers ja.

Han characters are not counted by that rule. A script in Chinese, which is written in Han characters, is not detected and falls to the English default. That is the trap: the job completes, you pay per character, and the result is not Mandarin. Japanese that is mostly kanji is also not guaranteed to be detected, so set ja there too.

Is Chinese a supported voice language on Sume?

Sume's Voices library tags voices with a fixed set of sixteen languages: en, ko, ja, zh, es, fr, de, pt, it, hi, nl, pl, ru, sv, tr and tl. Chinese is zh there, and the library has a sample sentence in Chinese for it. That tells you the product expects a voice to be labelled zh; it is not a promise about how any individual voice sounds.

The TTS reference itself names no supported-language count. Treat the answer as: ask for zh, pick a zh voice, listen. If a voice from another language is chosen, Sume returns 409 tts_voice_language_mismatch before creating a job, and you must confirm with confirm_language_mismatch: true to run it anyway.

A Mandarin request

Replace the voice id with a Chinese-tagged voice from your workspace.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: zh-intro-001" \
  -d '{
    "transcript": "大家好,这是我们新品的三十秒介绍。",
    "language": "zh",
    "voice": { "id": "'"$ZH_VOICE_ID"'" },
    "timestamps": { "words": true },
    "mode": "async"
  }'

What changes for timing and captions?

Chinese voiceover on Sume, read 2026-10-03
StepWhat the docs sayWhat to do for Chinese
SpeakSet language; omitted means EnglishSend zh on every request
Word timingstimestamps.words returns word start and end secondsCheck a sample: Chinese is not space-separated, so inspect how the units split
Burned captionsStyles are Latin faces or Hangul faces; no Chinese face is listedTest one clip before a batch, or ship a sidecar file where the platform takes one
BillingPer character, spaces and punctuation countCount your own characters; the rate applies to each one

What about Cantonese, Traditional characters and mixed text?

Sume's reference gives no separate language code for Cantonese or for Traditional versus Simplified characters. The Voices library tags voices with plain zh, and the TTS language field accepts a free string of 2 to 16 characters, so a code like zh-TW will pass the length check. Whether a given voice reads Traditional text or a regional variety differently is not documented, so the only answer is to test a line in the script you ship.

Mixed text is the other case to try. A Chinese sentence with an English brand name or a number in Latin digits is common in marketing copy. With language set to zh, listen for how the brand name is read; if it is wrong, TTS 1.0 accepts a pronunciation_dict_id for fixing a word without rewriting the sentence.

Finally, remember the language field is a hint for the speaker, not a translator. Sume does not translate the script for you in the TTS call, so the Chinese text must already be written by a translator you trust.

Keep the test cheap: a 40-character line costs well under a cent at $0.0475 per 1,000 characters, though a job never bills below the 1-cent floor. Run the same line with language omitted and with zh to hear the difference once, and keep both files as your team's reference for what the missing field sounds like.

Checklist before a Chinese batch

For the caption side see Burn Japanese, Chinese or Arabic captions. Pricing is $0.0475 per 1,000 characters for TTS 1.0 in the catalog, so a 300-character line is about 1.4 cents before any render.

  • Send language: "zh" explicitly, on every request, from a per-language config.
  • Use a voice tagged zh; confirm a mismatch only on purpose.
  • Listen to one line with numbers, a brand name and a date, since those are where reading goes wrong.
  • Do not burn Chinese captions with slam or punch untested: Sume documents no Chinese glyph face, and its Latin styles are not drawn for Han text.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume