HeyGen do_not_translate_terms vs Sume: handle terms before speech

HeyGen brand glossaries now take do_not_translate_terms and forced_translations for video translation. Sume has no glossary field; handle terms in your text.

4 min readSume
All posts

HeyGen's brand glossary endpoints now accept do_not_translate_terms and forced_translations, applied inside its video translation. The Sume docs read here list no glossary field, so with Sume you keep those rules in the text you translate before it is spoken.

HeyGen facts are from its API changelog; Sume facts from Audio detach, Timeline audio and Models, read 2026-10-01.

What did HeyGen add?

POST /v3/brand-glossaries and PATCH /v3/brand-glossaries/{brand_glossary_id} accept two lists beside terms. do_not_translate_terms keeps each term untranslated in every target language. forced_translations replaces a term with an exact translation, inserted verbatim. On PATCH, send a list to replace it, an empty list to clear it, or omit it to leave it alone. The changelog says the rules apply when a glossary is used by a translation feature, such as video translation over the API.

What does Sume give me for a translation workflow?

Sume audio surfaces from the docs, read 2026-10-01.
StepSurface
Extract the audioAudio detach: one workspace video in, a new audio artifact out
TranscribePOST /v1/stt-1.0/transcribe
Join audio for one renderTimeline 1.0 audio.parts[]
Talking shotFabric: audio_url, measured duration_seconds, exactly one visual source

Where do glossary rules go with Sume?

Into your own translation step. Keep a list of terms to leave alone and a map of forced translations, apply it to the transcript, and check the result before sending text to speech. That is a pattern, not a Sume feature. The pipeline is outlined in build an AI dubbing pipeline.

What should I check before moving a glossary over?

Export your terms and forced translations from HeyGen first; the changelog notes GET returns them whenever a glossary has rules. Then test the pronunciation of untranslated terms in each target language, since the speech step is where an untouched brand name can still be read wrongly.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume