Eleven v4 IPA support vs Sume's pronunciation_dict_id for brand names

ElevenLabs says Eleven v4 improves IPA phoneme support. Sume TTS has an optional pronunciation_dict_id. How each helps with brand names today.

4 min readSume
All posts

Brand names are where synthetic voices fail in public. ElevenLabs' Eleven v4 announcement says International Phonetic Alphabet (IPA) phoneme support is improved (read 2026-10-05), which lets you write how a word should sound instead of how it is spelled.

What Sume offers

Sume's TTS request schema has an optional pronunciation_dict_id, a string up to 128 characters described as a pronunciation dictionary id. The public API reference I read does not list an endpoint that creates a dictionary, so treat it as a field you can use only if you already hold an id. I did not find inline IPA markup in the request.

What works without either

  • Respell the word in the script: write "Soo-may" for a name the voice stumbles on.
  • Isolate it: put the problem sentence in its own TTS job and try two or three spellings.
  • Spell numbers and dates as words, as you want them said.
  • Set language correctly; a wrong language mangles every proper noun.

A cheap pronunciation test

Write the hard word in a 60-character sentence and render each respelling as its own job. Each is 1 cent, the minimum at $0.0475 per 1,000 characters, so five spellings cost 5 cents. Keep the winning spelling in a glossary file that your script tooling applies automatically.

Which to pick

If you need IPA control for many names and are happy on ElevenLabs, its tags are the direct tool. If your pipeline runs on Sume, keep a respelling glossary, and keep the check in the process: listen to the first line of every new script before you render all of it.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume