Chinese text to speech API: set language zh or it reads as English
Sume TTS only guesses Korean and Japanese when the language is missing. For Mandarin send language zh and pick a voice tagged zh, then test one line.

To make Mandarin speech with Sume's text to speech API, send the Chinese text as transcript, set language to zh, and choose a voice that is tagged for Chinese. Do not leave language out: Sume's fallback only recognises Hangul and kana, so a Chinese script with no language set is handed to an English-defaulted engine.
This follows the API reference, read 2026-10-03. Sume's reference does not publish a list of TTS languages, so what follows separates what is documented from what you should verify with one test line.
What does the fallback actually guess?
The language field is documented as the language the voice speaks the transcript in, with ko, ja and en as examples. Omitted, it defaults to English at the provider. Sume adds one rescue: if the text is mostly Hangul it infers ko, and if it is mostly kana it infers ja.
Han characters are not counted by that rule. A script in Chinese, which is written in Han characters, is not detected and falls to the English default. That is the trap: the job completes, you pay per character, and the result is not Mandarin. Japanese that is mostly kanji is also not guaranteed to be detected, so set ja there too.
Is Chinese a supported voice language on Sume?
Sume's Voices library tags voices with a fixed set of sixteen languages: en, ko, ja, zh, es, fr, de, pt, it, hi, nl, pl, ru, sv, tr and tl. Chinese is zh there, and the library has a sample sentence in Chinese for it. That tells you the product expects a voice to be labelled zh; it is not a promise about how any individual voice sounds.
The TTS reference itself names no supported-language count. Treat the answer as: ask for zh, pick a zh voice, listen. If a voice from another language is chosen, Sume returns 409 tts_voice_language_mismatch before creating a job, and you must confirm with confirm_language_mismatch: true to run it anyway.
A Mandarin request
Replace the voice id with a Chinese-tagged voice from your workspace.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: zh-intro-001" \
-d '{
"transcript": "大家好,这是我们新品的三十秒介绍。",
"language": "zh",
"voice": { "id": "'"$ZH_VOICE_ID"'" },
"timestamps": { "words": true },
"mode": "async"
}'What changes for timing and captions?
| Step | What the docs say | What to do for Chinese |
|---|---|---|
| Speak | Set language; omitted means English | Send zh on every request |
| Word timings | timestamps.words returns word start and end seconds | Check a sample: Chinese is not space-separated, so inspect how the units split |
| Burned captions | Styles are Latin faces or Hangul faces; no Chinese face is listed | Test one clip before a batch, or ship a sidecar file where the platform takes one |
| Billing | Per character, spaces and punctuation count | Count your own characters; the rate applies to each one |
What about Cantonese, Traditional characters and mixed text?
Sume's reference gives no separate language code for Cantonese or for Traditional versus Simplified characters. The Voices library tags voices with plain zh, and the TTS language field accepts a free string of 2 to 16 characters, so a code like zh-TW will pass the length check. Whether a given voice reads Traditional text or a regional variety differently is not documented, so the only answer is to test a line in the script you ship.
Mixed text is the other case to try. A Chinese sentence with an English brand name or a number in Latin digits is common in marketing copy. With language set to zh, listen for how the brand name is read; if it is wrong, TTS 1.0 accepts a pronunciation_dict_id for fixing a word without rewriting the sentence.
Finally, remember the language field is a hint for the speaker, not a translator. Sume does not translate the script for you in the TTS call, so the Chinese text must already be written by a translator you trust.
Keep the test cheap: a 40-character line costs well under a cent at $0.0475 per 1,000 characters, though a job never bills below the 1-cent floor. Run the same line with language omitted and with zh to hear the difference once, and keep both files as your team's reference for what the missing field sounds like.
Checklist before a Chinese batch
For the caption side see Burn Japanese, Chinese or Arabic captions. Pricing is $0.0475 per 1,000 characters for TTS 1.0 in the catalog, so a 300-character line is about 1.4 cents before any render.
- Send
language: "zh"explicitly, on every request, from a per-language config. - Use a voice tagged
zh; confirm a mismatch only on purpose. - Listen to one line with numbers, a brand name and a date, since those are where reading goes wrong.
- Do not burn Chinese captions with
slamorpunchuntested: Sume documents no Chinese glyph face, and its Latin styles are not drawn for Han text.
Sources
Related posts
More in Developers
- allowManagedModsOnly in Claude Code: does hosted Sume MCP still load?
allowManagedModsOnly keeps users' own Claude Code mods from loading. What it leaves alone, how a policy mod reviews the rest, and the Sume MCP connection.
- Claude Code mod: stop paid Sume calls after N in a session
Write a Claude Code mod that counts paid Sume MCP calls with a tool.call hook, denies call N+1, and fails closed. Code, matcher, and what it cannot cap.
- Claude Code mod hook skipped and a paid Sume call still ran
A Claude Code mod hook that throws or times out is skipped, so the call runs anyway. Add .catch to fail closed, and know what a mod still cannot guarantee.
- Claude Code mod: ask before a paid Sume call (and claude -p)
A Claude Code mod can hold a paid Sume MCP call with $.ui.ask and show max_spend_usd in the question. In claude -p nobody answers, so it refuses. Code inside.
Written by Sume