Brand name misheard in YouTube auto captions: burn your own script

YouTube warns that auto captions can misread accents and noisy audio. For a brand name that must be right, send script_text to Sume and burn your own wording.

5 min readSume
All posts

Short answer

If YouTube's automatic captions get your brand name wrong, burn captions you control. Send your exact wording as script_text to Sume's video captions API: Sume keeps the speech-to-text timings and aligns your text to them, so the spelling is yours and the timing follows the voice.

What YouTube says about accuracy

YouTube's page on automatic captions says they cover 80 or more languages and warns that accuracy can suffer with accents and background noise, and that generation can fail in cases such as extended silence at the start or overlapping speakers (read 2026-10-05). It also says potentially inappropriate words are masked in auto-generated captions only. Auto captions are a fine default and a poor guarantee for a product name.

What Sume does instead

With script_text, Sume does not use the transcript words as the text. It uses their timings as the source of truth for time and aligns the burned-in text to your script. If your script and the audio disagree too much, the job fails with a typed error, script_alignment_mismatch or script_alignment_failed, and the suggested next action is to simplify the script or omit it. A job that fails at alignment is better than a caption that silently says something else.

Auto captions and script_text compared, read 2026-10-05
QuestionYouTube auto captionsSume script_text
Who decides the spellingThe recognizerYou
Timing sourceThe recognizerSpeech-to-text timings
Known weak spotsAccents, noise, silent starts, overlapScript must align with the audio
Masked wordsMasked in auto-generated captionsBurned as you wrote them
Failure modeCaptions may not generateTyped alignment error

A request

The body sends the clip URL, a style and the script. The URL must be public HTTPS. The job costs $0.20 for a video up to 60 seconds. If the first look is wrong, restyle with source_caption_id and no second transcription, and send words only to correct the text.

import json, os, urllib.request

key = os.environ.get("SUME_API_KEY", "")
if not key:
    raise SystemExit("set SUME_API_KEY")
body = {
    "video_url": "https://media.sume.com/artifacts/artf_demo/clean.mp4",
    "style": "punch",
    "script_text": "Say hello to Zyntrex, the one-tap invoice app.",
}
req = urllib.request.Request(
    "https://api.sume.com/v1/video-captions",
    data=json.dumps(body).encode(),
    headers={"Authorization": "Bearer " + key,
             "Content-Type": "application/json",
             "Idempotency-Key": "caption-brand-001"},
)
print(json.load(urllib.request.urlopen(req)))

When to use which

Keep YouTube's captions on for accessibility and search, and burn your own for the hook line or any name that must be spelled right. A burned-in caption is pixels, so it cannot be turned off by the viewer; keep it to short, high-value text. The caption docs list the styles and the design overrides.

A practical workflow is two passes. On the first, run the clip with script_text and look at the result for the names that matter. If the alignment fails, shorten the script to the lines you need burned and drop filler, since the typed error recommends simplifying. On the second, if you want a different look, restyle with the same source_caption_id: the video and word timings are reused, the price stays at one caption job, and you avoid a second transcription.

Keep a glossary of the names that must never be wrong, such as product names, people and numbers, and paste the matching sentence into script_text for each clip. It costs nothing extra and turns a recurring correction task into a one-line input.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume