Sume TTS language field: set ja or ko, or the voice reads in English

Omit language on Sume TTS and the provider defaults to English; Sume infers ko or ja only from a Hangul- or kana-only script. Set it on non-English scripts.

5 min readSume
All posts

Set language on every non-English Sume TTS request. If you omit it, the provider defaults to English; Sume infers ko or ja only as a fallback when the transcript is all Hangul or all kana. A mixed script, such as Japanese text with a Latin brand name, is not covered by that fallback.

The field takes a BCP-47 or ISO-639 code of 2 to 16 characters, such as ko, ja or en. This is from the TTS request schema in the repository (read 2026-10-09).

What the schema says

The schema's guidance is blunt: never translate a non-English request into English, and set the language for every non-English transcript. A separate flag, confirm_language_mismatch, exists for the case where the voice and the language do not match. Send it only after the user confirms the warning, and omit it on the first request.

Language handling on Sume TTS, from the schema as of 2026-10-09
CaseWhat Sume does
language omitted, English textProvider default (English)
language omitted, Hangul-only or kana-only textSume infers ko or ja as a fallback
language omitted, mixed scriptNo inference; treated as English by the provider
language setUsed as the language the voice speaks the transcript in
voice and language mismatchA warning; resend with confirm_language_mismatch true after the user agrees

Why a default of English is expensive

A wrong-language read is billed like a right one: the charge is per character, so a 5,000-character Japanese script read as English costs 5 x $0.0475 = $0.2375 and gives you nothing to use. Multiply that by a batch of 200 scripts and you have paid $47.50 for audio you must regenerate.

The cheap defense is a one-line test per language and a request builder that refuses to send a non-English transcript without language. Make the field required in your own code even though the API treats it as optional.

A Korean example with the price

A 600-character Korean product line costs 0.6 x $0.0475 = $0.0285, and characters count the same way as Latin ones, including spaces and punctuation. The request is the same as an English one plus "language": "ko".

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ko-line-001" \
  -d '{"transcript": "안녕하세요. 오늘의 신제품을 소개합니다.",
       "avatar_handle": "studio_presenter", "language": "ko"}'

Captions follow the language too

If the audio goes under a video that you caption, the video captions job needs a Hangul style for Korean speech. slam, punch and tiktok-green use Latin faces and reject Hangul text with caption_hangul_text_latin_style. The language field on the caption job is only a speech-to-text hint and never picks the style.

A pre-flight checklist

Before a multilingual batch, check four things: every request carries language, the voice you picked speaks that language, mixed-script lines have been tested on their own, and the character counts include spaces and punctuation. A one-line test per language costs about a cent.

If the voice-language mismatch warning appears, do not set confirm_language_mismatch blindly. Ask whether the voice or the language is wrong, since confirming does not change the requested voice or language.

The same applies to the TTS Router: its request is TTS 1.0's plus a required model, and language works the same way there. Both are billed at $0.0475 per 1,000 characters.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume