How to sync subtitles with video: offset, drift, re-time
Sync subtitles by shifting every timestamp when all lines are off by the same amount, or scaling them when the gap grows. How to fix both for good.

To sync subtitles with a video, first check what kind of error you have. If every line is early or late by the same amount, add or subtract that offset from every timestamp. If the gap grows through the video, the subtitles were timed to a copy that plays at a different speed, so multiply every timestamp by one ratio. If only some lines are off, edit those lines, or re-time the whole file from the speech.
Fixing the .srt file itself makes the fix permanent in every player, unlike a player's delay setting. The frame rates below are FFmpeg's exact definitions from its utilities documentation, and the Sume facts come from the Video captions docs, both read on 2026-09-29. Anything called current behavior is read from Sume's code. If the voice itself is out of step with the picture, the subtitles are not the problem: see fix audio delay in a video.
How do I fix subtitles that are all early or late?
Find one line whose words you can hear clearly, note when it is spoken in the video and when the .srt shows it, and take the difference. A positive number delays the subtitles; a negative one brings them forward. The script below rewrites every HH:MM:SS,mmm timestamp in the file, scaling first and then shifting, so the same script fixes drift too.
import re, sys
# usage: python fix_srt.py in.srt out.srt OFFSET_SECONDS [RATIO]
src, dst, offset = sys.argv[1], sys.argv[2], float(sys.argv[3])
ratio = float(sys.argv[4]) if len(sys.argv) > 4 else 1.0
TIME = re.compile(r"(\d+):(\d\d):(\d\d),(\d{3})")
def fix(m):
h, mi, s, ms = map(int, m.groups())
t = max(0.0, (h * 3600 + mi * 60 + s + ms / 1000) * ratio + offset)
h, rest = divmod(round(t * 1000), 3_600_000)
mi, rest = divmod(rest, 60_000)
s, ms = divmod(rest, 1000)
return f"{h:02}:{mi:02}:{s:02},{ms:03}"
with open(src, encoding="utf-8-sig") as f:
text = f.read()
with open(dst, "w", encoding="utf-8") as f:
f.write(TIME.sub(fix, text))Why do subtitles drift further out of sync over time?
Because the subtitles were timed against a copy of the same film that runs at a different speed. A copy played at 25 frames per second instead of 23.976 shows each frame sooner, so a line spoken one hour into the 23.976 version comes 147 seconds earlier on that copy. Subtitles timed to it start close and end far off.
Measure two lines, one near the start and one near the end. The ratio is (video time of the last line − video time of the first) ÷ (subtitle time of the last − subtitle time of the first). The offset is the first line's video time minus (its subtitle time × the ratio). Pass both to the script. If the ratio lands near a value in the table, use the exact one.
| Subtitles timed at (fps) | Your video (fps) | Multiply times by |
|---|---|---|
25 (pal) | 23.976 (ntsc-film) | 1.0427 |
23.976 (ntsc-film) | 25 (pal) | 0.9590 |
25 (pal) | 24 (film) | 1.0417 |
24 (film) | 25 (pal) | 0.9600 |
24 (film) | 23.976 (ntsc-film) | 1.0010 |
23.976 (ntsc-film) | 24 (film) | 0.9990 |
Can I re-time subtitles from the speech automatically?
Yes, for burned-in subtitles on short clips. Sume's caption job (POST /v1/video-captions) takes a public HTTPS video_url. Send your subtitle text as script_text, and Sume keeps the speech-to-text word timings as the timing source and aligns your wording to them. The old file's timestamps are not used at all.
- Alignment can fail with
script_alignment_mismatchorscript_alignment_failed; the suggested next action is to simplifyscript_textor omit it. script_textis described for an English or Korean voiceover script.- If you already fixed the times with the script above, send the lines as
cues(text,start,endin seconds) instead: they skip speech-to-text and burn exactly that copy at those times. Sume takes no .srt upload; burn an SRT file into a video shows the conversion.
What doesn't this fix?
- It returns no corrected .srt. The caption job returns a captioned
video_url, and raw transcripts are not part of its public contract. For a soft subtitle file, keep the script's output. - In current code the caption job refuses a source over 60 seconds, or one with no audio stream even when you send cues. A clip with no audible speech fails as
caption_no_speechon thescript_textpath. - Subtitles that are wrong in only a few places need those lines edited by hand; edit auto-generated subtitles, then burn them covers that loop.
Sources
Related posts
More in Media tools
- TikTok TopView ad specs: size, length, and safe zones
TikTok TopView ads are vertical 9:16 videos of 540×960 px or more, 5–60 s (9–15 s recommended), up to 500 MB, 2,500 kbps or more, pre-approved.
- WooCommerce product image size: defaults and crop settings
WooCommerce shows product images 600 px wide uncropped, grid thumbnails 300 px wide cropped square, and gallery thumbnails at 100×100 by default.
- YouTube audio ads: specs, length, and the 2026 US change
YouTube audio ads run up to 30 seconds as a YouTube video, not an MP3. From Sept. 1, 2026, US-targeted ones sell only through Pandora and SiriusXM.
- YouTube TV ad specs: ad lengths and the 1-second rule
YouTube TV ads run :06 or shorter, :15 or shorter, :30, or :60, and Google says 30- and 60-second ads that miss by over 1 second won't serve.
Written by Sume