YouTube Auto-sync captions vs Sume script_text: track or burned in

YouTube Auto-sync times your transcript into a caption track. Sume script_text aligns your wording into burned-in captions. What each needs and which to use.

5 min readSume
All posts

YouTube's Auto-sync takes a transcript you type or upload, times it against the audio, and publishes it as a caption track on the video (read 2026-10-03). Sume does the same alignment idea for burned-in captions: pass script_text to POST /v1/video-captions and your wording replaces the speech-to-text wording on screen, timed to the speech. The two produce different things: a track the viewer can switch off, or pixels in the video.

Pick by where the video will be watched and whether the captions must survive a repost.

What does YouTube Auto-sync need?

On YouTube's Add subtitles page, Auto-sync is one of the ways to add a track in YouTube Studio. You enter the words in the video or upload a transcript file, and YouTube sets the timings, which can take a few minutes. The page lists the conditions: the transcript must be in a language its speech recognition supports, it must be the same language as the one spoken, and it is not recommended for videos over an hour long or with poor audio quality (read 2026-10-03).

Once it is ready, the result is published on the video as subtitles. The same page notes that "Type manually" also sets timings automatically, and suggests typing cues such as [applause] so viewers know what is happening.

YouTube Auto-sync against Sume script_text, read 2026-10-03
QuestionYouTube Auto-syncSume script_text
What you getA caption track on the videoA new video file with captions burned in
Viewer can switch offYes, it is a trackNo, the text is in the frames
Language ruleSame language as spoken, supported by YouTubeOptional language hint; omitted means automatic detection
InputTyped or uploaded transcriptscript_text string on the request
FailureNot recommended for poor audio or over an hourscript_alignment_mismatch or script_alignment_failed
Silent clipsNot applicable, needs speechFails as caption_no_speech; use cues instead

How does Sume script_text work?

With script_text, Sume keeps speech-to-text word timings as the timing source and aligns the burned-in wording to your script (video captions). So the screen shows your spelling, product names and punctuation, while the timing still comes from the audio. If the script differs too much from what is said, the job fails with script_alignment_mismatch or script_alignment_failed, and the suggested next action is to simplify the script or omit it. The failure modes are covered in the mismatch fix post.

Note the limit: the public resource returns the captioned video and artifacts, not a transcript or timing file, and SRT uploads are unsupported. So Sume cannot hand you an aligned SRT to attach to YouTube. For a track on YouTube, use Auto-sync there.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: script-burn-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": "slam",
    "language": "en",
    "script_text": "Say hello to the Sume developer platform."
  }'

Which should I use?

The choice follows the destination, not the tool.

  • Long-form on YouTube, where viewers expect a toggle: upload to Studio and use Auto-sync, or upload a file with timing.
  • Vertical clips that will be downloaded and reposted to other apps: burn with Sume so the captions travel with the file.
  • A clip with unusual brand spellings: script_text fixes them in the burned text; on YouTube, edit the track after Auto-sync.
  • A silent clip: neither alignment path works, since both need speech; send cues with text, start and end.
  • Both: burn short-form versions with Sume and keep the YouTube track for the full upload.

What goes wrong in practice?

Both paths fail for the same underlying reason: the script and the audio disagree. A transcript that was written before the take, with lines the speaker skipped, reworded or added, gives any aligner nothing to hold on to. On YouTube the symptom is a slow or wrong sync, and the page itself steers you away from poor audio. On Sume the symptom is a typed error, which is easier to automate around: catch script_alignment_mismatch, then retry once without script_text so the speech-to-text wording is burned instead.

The cheapest fix is upstream. Record from the script, or write the script from the transcript you actually have, and keep punctuation and numbers as spoken. If a line must differ on screen from what is said, say a legal disclaimer, put it in a separate overlay with cues rather than forcing it through alignment, because script_text and cues are mutually exclusive on one request.

What does it cost?

YouTube's page lists no price for Auto-sync. A Sume caption job is $0.20 per job for videos up to 60 seconds under the current fixed estimate; confirm live pricing in GET /v1/catalog. If the text was not written for the video, transcribe first and burn after you edit it, which avoids paying for a failed alignment. If you want to restyle later without a second transcription, see the restyle post.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume