Sume TTS 400 tts_language_script_mismatch: Korean encoding check

Sume TTS rejects language ko when the transcript has no Hangul syllable: 400 tts_language_script_mismatch, no charge. Usually the text was decoded wrongly.

5 min readSume
All posts

If you send language: "ko" to Sume TTS with a transcript that has no Hangul syllable, the API returns 400 tts_language_script_mismatch with details.field set to transcript. The check runs before a job is queued, so nothing is reserved, billed or sent to a voice provider; the usual cause is text that was decoded with the wrong character set before it reached you.

The behaviour is in the API source and its tests in the repository, and the language field is in the Sume OpenAPI contract. This post contains no Korean text on purpose, so the examples describe the bad inputs instead of showing them.

What is checked

The rule is narrow. When language is ko (case-insensitive) and the transcript contains no character in the Hangul syllable block, the request fails with a message that says Korean TTS needs at least one Hangul syllable and tells you to check the original text and its encoding. Hangul letters on their own, without a composed syllable, do not satisfy it, and neither do digits, punctuation or Latin text. The same check runs on the TTS 1.0 and TTS router routes.

Inputs with language ko (per the API tests in the repository, read 2026-10-10)
TranscriptResult
Valid Korean with Hangul syllablesAccepted
Korean bytes decoded as Latin-1 (mojibake)400 tts_language_script_mismatch
An English sentence400 tts_language_script_mismatch
Digits and punctuation only400 tts_language_script_mismatch
Isolated Hangul letters only, no syllables400 tts_language_script_mismatch

Why mojibake is the usual cause

The tests name the production failure directly: UTF-8 Korean decoded as Latin-1. That produces a string of accented Latin characters that looks like text to a program but contains no Hangul at all. Without this guard, a voice would be asked to read garbage and the user would be charged for it. Typical places for the mistake are a CSV opened with the wrong encoding, a shell that does not use UTF-8, or an HTTP client that sets a charset incorrectly.

How to fix it

Check the text as your code holds it, before the request:

  • Print the transcript's code points rather than the characters.
  • Make sure the file is read as UTF-8 and the request body is sent as UTF-8 JSON.
  • If the text really is English, remove language: "ko" or set language to en.
  • Do not retry the same body. The error is marked non-retryable and the action is to fix the input.

When language is omitted

The OpenAPI description says an omitted language defaults to English at the provider, but that Sume infers ko or ja from a transcript made only of Hangul or kana as a fallback. So the safest habit is to set language explicitly on every non-English request and let this check protect you.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume