Text to speech pronunciation: how to fix wrong words

Fix text to speech pronunciation: set the text's language, use a voice recorded in it, and respell tricky words. What Sume's TTS API checks for you.

5 min readSume
All posts

To fix text to speech pronunciation, start with the input: tell the engine which language the text is in, use a voice recorded in that language, and write names, numbers, and abbreviations the way they should be said. Then regenerate one short line and listen before you redo the whole script. With Sume's TTS API that means sending language with every non-English request, since an omitted language defaults to English, and taking its 409 voice-language double-check seriously.

The Sume facts come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28. Checks called current behavior are read from Sume's code, and the price from the code behind API pricing. The respelling tips are general techniques to try, not documented Sume features.

Why does text to speech mispronounce words?

Look for three causes. The engine reads the text with the rules of the wrong language. The voice was recorded in a different language from the script. Or the word itself can be read more than one way: names, brand names, acronyms, numbers, dates, and abbreviations.

Sume's schema describes language as the language the voice speaks the transcript in, and says to set it for every non-English transcript: omitted, it defaults to English, with Korean or Japanese inferred only from a Hangul- or kana-only transcript as a fallback. So a French or Spanish script sent without its code is synthesized with the English default.

How do I make the engine use the right language?

  • Send language as a BCP-47 / ISO-639 code, such as ko, ja, or en. The field takes one code, so a script that switches languages is split into one request per language, as in multilingual text to speech.
  • For Korean, write the script in Hangul. In current code, language: "ko" with no Hangul syllable fails with 400 tts_language_script_mismatch, so romanized Korean is refused; see text to speech in Korean.
  • For Japanese, write kana and kanji and send ja; in current code the fallback never infers Japanese from kanji alone, as text to speech in Japanese explains.

Why did I get a 409 saying pronunciation may sound unnatural?

Because the voice and the script disagree on the language. In current code, when a voice from Sume's Voices library has a recorded language that differs from the request's, the submit stops with 409 tts_voice_language_mismatch. The message names both languages and warns: “Pronunciation may sound unnatural.” No job or charge exists at that point.

Pick a voice recorded in the script's language, or resend the same request and Idempotency-Key with confirm_language_mismatch: true. Confirming does not change the voice or the language; it only accepts the risk. AI voiceover in your own voice shows the same check for a cloned voice.

Pronunciation-related fields and checks, from the TTS schema in the Sume API reference and Sume's current code, read 2026-09-28.
Field or checkWhat it does
languageThe language the voice speaks the transcript in. Omitted: English, except the Hangul- or kana-only fallback.
Voice's recorded languageLibrary voices record one of 16 languages; a mismatch answers 409 tts_voice_language_mismatch before any job (current code).
confirm_language_mismatchtrue after you accept the warning. Changes neither the voice nor the language.
language: "ko" without Hangul400 tts_language_script_mismatch (current code).
pronunciation_dict_idAn optional pronunciation dictionary id. No route or format for creating a dictionary is documented.

How do I respell a word so it's read correctly?

The voice speaks the transcript as written, so the lever you control is the spelling you send. None of these is a documented Sume feature; they are rewrites to try on a single test sentence:

  • Write numbers, dates, prices, and symbols out as words: “twenty twenty-six”, “the fourth of July”, “five dollars”.
  • For an acronym read letter by letter, separate the letters (“S Q L”). For one read as a word, write the word (“sequel”).
  • Respell a name or borrowed word the way it sounds, then keep the respelling in a lookup table your code applies before every request, so each script gets the same fix.
  • Put the word in a short sentence rather than alone, and listen to the result before you regenerate a long script.

Does Sume support SSML or a pronunciation dictionary?

Not as a documented feature. The schema describes transcript only as text to synthesize and documents no SSML or phoneme markup. It does accept pronunciation_dict_id, described as an optional pronunciation dictionary id, but the API reference documents no way to create or format a dictionary, so don't build on it.

What does testing a pronunciation fix cost?

Every try is a new TTS job billed on its transcript characters at $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, spaces and punctuation included. Test the fix on one short sentence, then send the full script once the word sounds right.

Sources

Related posts

More in Models

All Models posts

Written by Sume