HeyGen brand glossary: speech changes, captions keep spelling

HeyGen's brand_glossary_id changes synthesized speech only; captions keep the original spelling. Sume has TTS pronunciation_dict_id and a caption stage.

4 min readSume
All posts

In HeyGen, a brand glossary with do_not_translate_terms and forced_translations is applied through brand_glossary_id, and the changelog says pronunciation affects the synthesized audio only: captions and subtitles keep the original spelling. On Sume the same split exists in a different shape: TTS takes an optional pronunciation_dict_id, and captions are their own stage.

HeyGen facts are from its changelog; Sume facts from the OpenAPI schema and Avatar videos, read 2026-10-01.

Why does the HeyGen changelog stress speech only?

Because a glossary entry fixes how a term sounds, not how it is written. When POST /v3/video-agents takes a brand_glossary_id, the narration it writes and voices pronounces your terms your way, while captions and subtitles stay as the original spelling. The glossary is set at session creation and covers the whole session. That means a team that wants a respelled brand name on screen has to handle the caption text itself.

What is the Sume equivalent?

The TTS request has an optional pronunciation_dict_id (a string up to 128 characters). The sources read here do not describe do-not-translate or forced-translation rules, so for translation, keep your term list in your own pipeline and apply it before you send text to TTS. A speech language hint also exists on speech-to-text: an optional BCP-47 value such as en or ko, omitted for auto-detect.

Where do captions fit?

Brand-term handling in speech and captions, from the HeyGen changelog and the Sume docs, read 2026-10-01.
SurfaceHeyGenSume
Speechbrand_glossary_id changes pronunciationpronunciation_dict_id on TTS
CaptionsKeep the original spellingSeparate caption stage; Korean needs a Hangul style
Caption failureNot describedSoft-fails; the video can still succeed

How do I keep spelling right in captions?

Treat captions as text you supply, not as a readback of the audio. For Korean speech, pick a Hangul caption style, because a Latin style is rejected. If the caption stage fails, the avatar job can still succeed with a clean primary video_url and captions.status=failed, so check that field before publishing. The spelling workflow is in captions with brand name spelling.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume