Forgot the TTS language field? Sume only infers Korean and Japanese

Omit language on a Sume TTS request and the voice check assumes English. Sume infers ko or ja only when Hangul or kana outnumber Latin letters.

3 min readSume
All posts

A common bug in a TTS integration: the script is Spanish, the voice is Spanish, and the audio has an English accent. The usual cause is a missing language field.

What Sume does

Without language, the provider reads the text as English. Sume adds one inference, based on letter counts in the transcript: Hangul is sent as Korean when it outnumbers Latin letters, and kana as Japanese when it outnumbers both Hangul and Latin letters. Anything else, including Spanish, French, German and Chinese with Han characters, stays English unless you say otherwise.

Fix

  • Always send language as a short code such as es, fr, de, pt.
  • Use a voice tagged for that language in the voice library.
  • With no language, a Spanish text defaults to English for the voice check, so a Spanish voice gets a 409 tts_voice_language_mismatch before any job is queued or charge made. An English voice does not trigger it, which is how the English-accent bug gets through.
  • Keep a unit test that sends one Spanish sentence and asserts the field is present in your request body.

Edge cases

A script that mixes Korean and English is inferred by counts, so a line with more Latin letters than Hangul stays English. Set ko yourself, or split the script per language and join the audio with Timeline concat.

Related posts

More in Developers

All Developers posts

Written by Sume