YouTube Auto-sync captions vs Sume script_text: track or burned in
YouTube Auto-sync times your transcript into a caption track. Sume script_text aligns your wording into burned-in captions. What each needs and which to use.

YouTube's Auto-sync takes a transcript you type or upload, times it against the audio, and publishes it as a caption track on the video (read 2026-10-03). Sume does the same alignment idea for burned-in captions: pass script_text to POST /v1/video-captions and your wording replaces the speech-to-text wording on screen, timed to the speech. The two produce different things: a track the viewer can switch off, or pixels in the video.
Pick by where the video will be watched and whether the captions must survive a repost.
What does YouTube Auto-sync need?
On YouTube's Add subtitles page, Auto-sync is one of the ways to add a track in YouTube Studio. You enter the words in the video or upload a transcript file, and YouTube sets the timings, which can take a few minutes. The page lists the conditions: the transcript must be in a language its speech recognition supports, it must be the same language as the one spoken, and it is not recommended for videos over an hour long or with poor audio quality (read 2026-10-03).
Once it is ready, the result is published on the video as subtitles. The same page notes that "Type manually" also sets timings automatically, and suggests typing cues such as [applause] so viewers know what is happening.
| Question | YouTube Auto-sync | Sume script_text |
|---|---|---|
| What you get | A caption track on the video | A new video file with captions burned in |
| Viewer can switch off | Yes, it is a track | No, the text is in the frames |
| Language rule | Same language as spoken, supported by YouTube | Optional language hint; omitted means automatic detection |
| Input | Typed or uploaded transcript | script_text string on the request |
| Failure | Not recommended for poor audio or over an hour | script_alignment_mismatch or script_alignment_failed |
| Silent clips | Not applicable, needs speech | Fails as caption_no_speech; use cues instead |
How does Sume script_text work?
With script_text, Sume keeps speech-to-text word timings as the timing source and aligns the burned-in wording to your script (video captions). So the screen shows your spelling, product names and punctuation, while the timing still comes from the audio. If the script differs too much from what is said, the job fails with script_alignment_mismatch or script_alignment_failed, and the suggested next action is to simplify the script or omit it. The failure modes are covered in the mismatch fix post.
Note the limit: the public resource returns the captioned video and artifacts, not a transcript or timing file, and SRT uploads are unsupported. So Sume cannot hand you an aligned SRT to attach to YouTube. For a track on YouTube, use Auto-sync there.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: script-burn-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "slam",
"language": "en",
"script_text": "Say hello to the Sume developer platform."
}'Which should I use?
The choice follows the destination, not the tool.
- Long-form on YouTube, where viewers expect a toggle: upload to Studio and use Auto-sync, or upload a file with timing.
- Vertical clips that will be downloaded and reposted to other apps: burn with Sume so the captions travel with the file.
- A clip with unusual brand spellings:
script_textfixes them in the burned text; on YouTube, edit the track after Auto-sync. - A silent clip: neither alignment path works, since both need speech; send
cueswith text, start and end. - Both: burn short-form versions with Sume and keep the YouTube track for the full upload.
What goes wrong in practice?
Both paths fail for the same underlying reason: the script and the audio disagree. A transcript that was written before the take, with lines the speaker skipped, reworded or added, gives any aligner nothing to hold on to. On YouTube the symptom is a slow or wrong sync, and the page itself steers you away from poor audio. On Sume the symptom is a typed error, which is easier to automate around: catch script_alignment_mismatch, then retry once without script_text so the speech-to-text wording is burned instead.
The cheapest fix is upstream. Record from the script, or write the script from the transcript you actually have, and keep punctuation and numbers as spoken. If a line must differ on screen from what is said, say a legal disclaimer, put it in a separate overlay with cues rather than forcing it through alignment, because script_text and cues are mutually exclusive on one request.
What does it cost?
YouTube's page lists no price for Auto-sync. A Sume caption job is $0.20 per job for videos up to 60 seconds under the current fixed estimate; confirm live pricing in GET /v1/catalog. If the text was not written for the video, transcribe first and burn after you edit it, which avoids paying for a failed alignment. If you want to restyle later without a second transcription, see the restyle post.
Sources
Related posts
More in Media tools
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
- How to extract frames from a video with the Sume API
POST /v1/video-frames returns stills at the times you name from one Sume-hosted clip, as durable images at source size. The call is unbilled.
- How to use Sume's Timeline compose and Timeline audio APIs
Timeline compose puts one still and one video in the same frame as a new MP4. Timeline audio joins or splits Sume-hosted audio into reusable files.
Written by Sume