How to sync subtitles with video: offset, drift, re-time

Sync subtitles by shifting every timestamp when all lines are off by the same amount, or scaling them when the gap grows. How to fix both for good.

5 min readSume
All posts

To sync subtitles with a video, first check what kind of error you have. If every line is early or late by the same amount, add or subtract that offset from every timestamp. If the gap grows through the video, the subtitles were timed to a copy that plays at a different speed, so multiply every timestamp by one ratio. If only some lines are off, edit those lines, or re-time the whole file from the speech.

Fixing the .srt file itself makes the fix permanent in every player, unlike a player's delay setting. The frame rates below are FFmpeg's exact definitions from its utilities documentation, and the Sume facts come from the Video captions docs, both read on 2026-09-29. Anything called current behavior is read from Sume's code. If the voice itself is out of step with the picture, the subtitles are not the problem: see fix audio delay in a video.

How do I fix subtitles that are all early or late?

Find one line whose words you can hear clearly, note when it is spoken in the video and when the .srt shows it, and take the difference. A positive number delays the subtitles; a negative one brings them forward. The script below rewrites every HH:MM:SS,mmm timestamp in the file, scaling first and then shifting, so the same script fixes drift too.

import re, sys

# usage: python fix_srt.py in.srt out.srt OFFSET_SECONDS [RATIO]
src, dst, offset = sys.argv[1], sys.argv[2], float(sys.argv[3])
ratio = float(sys.argv[4]) if len(sys.argv) > 4 else 1.0
TIME = re.compile(r"(\d+):(\d\d):(\d\d),(\d{3})")

def fix(m):
    h, mi, s, ms = map(int, m.groups())
    t = max(0.0, (h * 3600 + mi * 60 + s + ms / 1000) * ratio + offset)
    h, rest = divmod(round(t * 1000), 3_600_000)
    mi, rest = divmod(rest, 60_000)
    s, ms = divmod(rest, 1000)
    return f"{h:02}:{mi:02}:{s:02},{ms:03}"

with open(src, encoding="utf-8-sig") as f:
    text = f.read()
with open(dst, "w", encoding="utf-8") as f:
    f.write(TIME.sub(fix, text))

Why do subtitles drift further out of sync over time?

Because the subtitles were timed against a copy of the same film that runs at a different speed. A copy played at 25 frames per second instead of 23.976 shows each frame sooner, so a line spoken one hour into the 23.976 version comes 147 seconds earlier on that copy. Subtitles timed to it start close and end far off.

Measure two lines, one near the start and one near the end. The ratio is (video time of the last line − video time of the first) ÷ (subtitle time of the last − subtitle time of the first). The offset is the first line's video time minus (its subtitle time × the ratio). Pass both to the script. If the ratio lands near a value in the table, use the exact one.

Rates from FFmpeg's video rate abbreviations, read 2026-09-29. Ratios are computed from the exact rates (23.976 is 24000/1001).
Subtitles timed at (fps)Your video (fps)Multiply times by
25 (pal)23.976 (ntsc-film)1.0427
23.976 (ntsc-film)25 (pal)0.9590
25 (pal)24 (film)1.0417
24 (film)25 (pal)0.9600
24 (film)23.976 (ntsc-film)1.0010
23.976 (ntsc-film)24 (film)0.9990

Can I re-time subtitles from the speech automatically?

Yes, for burned-in subtitles on short clips. Sume's caption job (POST /v1/video-captions) takes a public HTTPS video_url. Send your subtitle text as script_text, and Sume keeps the speech-to-text word timings as the timing source and aligns your wording to them. The old file's timestamps are not used at all.

  • Alignment can fail with script_alignment_mismatch or script_alignment_failed; the suggested next action is to simplify script_text or omit it.
  • script_text is described for an English or Korean voiceover script.
  • If you already fixed the times with the script above, send the lines as cues (text, start, end in seconds) instead: they skip speech-to-text and burn exactly that copy at those times. Sume takes no .srt upload; burn an SRT file into a video shows the conversion.

What doesn't this fix?

  • It returns no corrected .srt. The caption job returns a captioned video_url, and raw transcripts are not part of its public contract. For a soft subtitle file, keep the script's output.
  • In current code the caption job refuses a source over 60 seconds, or one with no audio stream even when you send cues. A clip with no audible speech fails as caption_no_speech on the script_text path.
  • Subtitles that are wrong in only a few places need those lines edited by hand; edit auto-generated subtitles, then burn them covers that loop.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume