HeyGen brand glossary: speech changes, captions keep spelling
HeyGen's brand_glossary_id changes synthesized speech only; captions keep the original spelling. Sume has TTS pronunciation_dict_id and a caption stage.

In HeyGen, a brand glossary with do_not_translate_terms and forced_translations is applied through brand_glossary_id, and the changelog says pronunciation affects the synthesized audio only: captions and subtitles keep the original spelling. On Sume the same split exists in a different shape: TTS takes an optional pronunciation_dict_id, and captions are their own stage.
HeyGen facts are from its changelog; Sume facts from the OpenAPI schema and Avatar videos, read 2026-10-01.
Why does the HeyGen changelog stress speech only?
Because a glossary entry fixes how a term sounds, not how it is written. When POST /v3/video-agents takes a brand_glossary_id, the narration it writes and voices pronounces your terms your way, while captions and subtitles stay as the original spelling. The glossary is set at session creation and covers the whole session. That means a team that wants a respelled brand name on screen has to handle the caption text itself.
What is the Sume equivalent?
The TTS request has an optional pronunciation_dict_id (a string up to 128 characters). The sources read here do not describe do-not-translate or forced-translation rules, so for translation, keep your term list in your own pipeline and apply it before you send text to TTS. A speech language hint also exists on speech-to-text: an optional BCP-47 value such as en or ko, omitted for auto-detect.
Where do captions fit?
| Surface | HeyGen | Sume |
|---|---|---|
| Speech | brand_glossary_id changes pronunciation | pronunciation_dict_id on TTS |
| Captions | Keep the original spelling | Separate caption stage; Korean needs a Hangul style |
| Caption failure | Not described | Soft-fails; the video can still succeed |
How do I keep spelling right in captions?
Treat captions as text you supply, not as a readback of the audio. For Korean speech, pick a Hangul caption style, because a Latin style is rejected. If the caption stage fails, the avatar job can still succeed with a clean primary video_url and captions.status=failed, so check that field before publishing. The spelling workflow is in captions with brand name spelling.
Sources
Related posts
More in Developers
- HeyGen brand kit from a website URL vs Sume product and scene
HeyGen builds a stored brand kit from a website URL with POST /v3/brand-kits. Sume has no kit object: you pass product, scene and avatar inputs per request.
- HeyGen download_failed error: URL rules vs Sume's media errors
HeyGen download_failed means a media URL was not public, mismatched or corrupt. Sume returns image_not_fetchable or input_media_unreachable for similar cases.
- HeyGen get video scenes API vs Sume video inspect stills
HeyGen's scenes endpoint returns a typed composition. Sume's video inspect returns probe facts and stills for one hosted clip, not typed scenes.
- HeyGen image to video from a photo API vs Sume image_url
HeyGen type image animates a PNG or JPEG with a script and voice_id, no avatar setup. Sume does it with image_url plus audio on the Fabric route.
Written by Sume