Text to speech pronunciation: how to fix wrong words
Fix text to speech pronunciation: set the text's language, use a voice recorded in it, and respell tricky words. What Sume's TTS API checks for you.

To fix text to speech pronunciation, start with the input: tell the engine which language the text is in, use a voice recorded in that language, and write names, numbers, and abbreviations the way they should be said. Then regenerate one short line and listen before you redo the whole script. With Sume's TTS API that means sending language with every non-English request, since an omitted language defaults to English, and taking its 409 voice-language double-check seriously.
The Sume facts come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28. Checks called current behavior are read from Sume's code, and the price from the code behind API pricing. The respelling tips are general techniques to try, not documented Sume features.
Why does text to speech mispronounce words?
Look for three causes. The engine reads the text with the rules of the wrong language. The voice was recorded in a different language from the script. Or the word itself can be read more than one way: names, brand names, acronyms, numbers, dates, and abbreviations.
Sume's schema describes language as the language the voice speaks the transcript in, and says to set it for every non-English transcript: omitted, it defaults to English, with Korean or Japanese inferred only from a Hangul- or kana-only transcript as a fallback. So a French or Spanish script sent without its code is synthesized with the English default.
How do I make the engine use the right language?
- Send
languageas a BCP-47 / ISO-639 code, such asko,ja, oren. The field takes one code, so a script that switches languages is split into one request per language, as in multilingual text to speech. - For Korean, write the script in Hangul. In current code,
language: "ko"with no Hangul syllable fails with400 tts_language_script_mismatch, so romanized Korean is refused; see text to speech in Korean. - For Japanese, write kana and kanji and send
ja; in current code the fallback never infers Japanese from kanji alone, as text to speech in Japanese explains.
Why did I get a 409 saying pronunciation may sound unnatural?
Because the voice and the script disagree on the language. In current code, when a voice from Sume's Voices library has a recorded language that differs from the request's, the submit stops with 409 tts_voice_language_mismatch. The message names both languages and warns: “Pronunciation may sound unnatural.” No job or charge exists at that point.
Pick a voice recorded in the script's language, or resend the same request and Idempotency-Key with confirm_language_mismatch: true. Confirming does not change the voice or the language; it only accepts the risk. AI voiceover in your own voice shows the same check for a cloned voice.
| Field or check | What it does |
|---|---|
language | The language the voice speaks the transcript in. Omitted: English, except the Hangul- or kana-only fallback. |
| Voice's recorded language | Library voices record one of 16 languages; a mismatch answers 409 tts_voice_language_mismatch before any job (current code). |
confirm_language_mismatch | true after you accept the warning. Changes neither the voice nor the language. |
language: "ko" without Hangul | 400 tts_language_script_mismatch (current code). |
pronunciation_dict_id | An optional pronunciation dictionary id. No route or format for creating a dictionary is documented. |
How do I respell a word so it's read correctly?
The voice speaks the transcript as written, so the lever you control is the spelling you send. None of these is a documented Sume feature; they are rewrites to try on a single test sentence:
- Write numbers, dates, prices, and symbols out as words: “twenty twenty-six”, “the fourth of July”, “five dollars”.
- For an acronym read letter by letter, separate the letters (“S Q L”). For one read as a word, write the word (“sequel”).
- Respell a name or borrowed word the way it sounds, then keep the respelling in a lookup table your code applies before every request, so each script gets the same fix.
- Put the word in a short sentence rather than alone, and listen to the result before you regenerate a long script.
Does Sume support SSML or a pronunciation dictionary?
Not as a documented feature. The schema describes transcript only as text to synthesize and documents no SSML or phoneme markup. It does accept pronunciation_dict_id, described as an optional pronunciation dictionary id, but the API reference documents no way to create or format a dictionary, so don't build on it.
What does testing a pronunciation fix cost?
Every try is a new TTS job billed on its transcript characters at $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, spaces and punctuation included. Test the fix on one short sentence, then send the full script once the word sounds right.
Sources
Related posts
More in Models
- Video to anime AI: how to turn a clip into anime
To turn a video into anime with AI, use a video-to-video edit that names the style and what to keep, or restyle one frame and animate it.
- Wan 3.0 vs Seedance 2.5: 30-second clips, inputs and price
Wan 3.0 and Seedance 2.5 both make clips of up to 30 s with audio, frames and references. Wan starts at 2 s and bills per second, Seedance per token.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
Written by Sume