Keep brand names intact in translated, burned-in captions

Index-Translate supports glossary instructions, but vendor notes say compliance is not guaranteed. Check every term in code before Sume burns the cues in.

5 min readSume
All posts

To keep brand names unchanged in translated subtitles, tell the translation model which terms to keep, then verify each term in code before you burn the captions. Index-Translate lists glossary customization as a feature, but the model card for its speech sibling, Index-Echo, says context and glossary instructions do not guarantee exact compliance. A string check costs nothing and catches the misses before a render does.

Sume burns exactly the text you give it. Once the cues are final there is no second translation pass, so a wrong brand spelling in the cue is a wrong spelling on screen.

What does Index-Translate offer for terminology?

The README describes three instruction types on top of plain translation: a terminology glossary treated as a hard constraint, format preservation for structures such as JSON and CSV, and softer tone and genre controls. The inference scripts expose these as options; the glossary option is -g. The README's exact input format for the hosted API is not spelled out on the pages we read, so the safest route is to put the glossary in the prompt in plain words and test it.

Instruction types listed by the vendor, read 2026-10-04
InstructionVendor descriptionWhat to do on your side
Terminology glossaryHard constraint on named termsAdd a term list to the prompt, then check the output
Format preservationKeeps JSON, CSV or HTML structureUseful for sending cues as JSON
Tone and genreSoft constraintsReview a sample before a full run

How do you check the terms?

Keep a list of protected strings with the cues they belong to. After translation, confirm each protected string appears verbatim in the translated cue, and queue the failures for a retry or a manual fix. The helper below returns the cues that lost a term.

PROTECTED = ["Sume", "Avatar 1.0", "Index-Translate"]

def lost_terms(source_cues, translated_cues):
    problems = []
    for i, (src, out) in enumerate(zip(source_cues, translated_cues)):
        for term in PROTECTED:
            if term in src["text"] and term not in out["text"]:
                problems.append((i, term))
    return problems

source = [{"text": "用 Sume 生成 Avatar 1.0 视频", "start": 0, "end": 3}]
translated = [{"text": "Make a video with the avatar tool", "start": 0, "end": 3}]
print(lost_terms(source, translated))

What if you would rather not translate the cue at all?

If the spoken audio already carries the brand name and you only need it spelled correctly, Sume's script_text option takes a different route. It keeps the speech-to-text word timings and aligns the burned wording to your script, so your spelling wins. That is a same-language fix, not a translation, and it cannot be combined with cues.

For the translated route, send the checked cues to video captions. Each cue is limited to 400 characters, a request to 200 cues and 60 seconds, and a standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate.

Where do glossary misses usually happen?

Log the failed indices, retry those cues with the term repeated in the instruction, and fix any that still fail by hand. Re-burning a clip costs another job, so finish the checks first.

  • Names split across two cues, so neither half matches the term.
  • Terms that the model transliterates into the target script.
  • Plural or inflected forms, where an exact match fails even though the line is fine.
  • Short lines, where a protected term is most of the text.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume