YouTube .sbv vs .srt: the timestamp difference and a Python converter

YouTube accepts .sbv and .srt as basic caption files. They differ in time format and numbering. A short Python converter turns timed cues into either file.

4 min readSume
All posts

An .sbv file is YouTube's SubViewer format. It carries the same text as an .srt file, but each cue is a single line of times, 0:00:00.599,0:00:04.160, with no index number and a period before the milliseconds, where SubRip uses an index line and 00:00:00,599 --> 00:00:04,160, with a comma. YouTube's help page lists both as basic formats: plain UTF-8, no style markup recognized. If you hold timed cues, a dozen lines of Python write either one.

Everything about the formats below comes from YouTube's supported subtitle and closed caption files page, read 2026-10-03.

What is the difference between SubRip and SubViewer?

YouTube says the main difference is the format of the caption start and stop times. Its own examples show the pair side by side, and both accept speaker labels such as >> ALICE: and sound cues such as [intro music] as ordinary text.

YouTube Help, Supported subtitle and closed caption files, read 2026-10-03.
FormatExtensionCue layout in YouTube's exampleStyle markup
SubRip.srtIndex line, then 00:00:00,599 --> 00:00:04,160, then textNot recognized; plain UTF-8 only
SubViewer.sbv or .sub0:00:00.599,0:00:04.160, then text; no index lineNot recognized; plain UTF-8 only
WebVTT.vttNot shown in the page's examplesPositioning supported; styling limited to <b>, <i>, <u>
TTML.ttmlNot shown in the page's examplesStyling and positioning supported

Where do the timed cues come from?

If the audio is already transcribed, you have them. Sume's video inspect route runs speech-to-text when you send transcribe: true, at $0.01 per audio minute, and segmentation.mode: "sentence" returns gapless sentence segments shaped like caption lines. The converter below expects each cue as text plus start and end in seconds, the same three fields a Sume caption cue uses, so map your segments to that shape first and check the field names against a real response.

What does the converter look like?

The function splits seconds into hours, minutes, seconds and milliseconds, then formats them for each file. SubRip pads the hour to two digits and uses a comma; SubViewer leaves the hour unpadded and uses a period.

def parts(t):
    ms = round(t * 1000)
    return ms // 3600000, ms // 60000 % 60, ms // 1000 % 60, ms % 1000

def srt_time(t):
    h, m, s, ms = parts(t)
    return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}"

def sbv_time(t):
    h, m, s, ms = parts(t)
    return f"{h}:{m:02d}:{s:02d}.{ms:03d}"

def to_srt(cues):
    return "\n".join(
        f"{i}\n{srt_time(c['start'])} --> {srt_time(c['end'])}\n{c['text']}\n"
        for i, c in enumerate(cues, 1))

def to_sbv(cues):
    return "\n".join(
        f"{sbv_time(c['start'])},{sbv_time(c['end'])}\n{c['text']}\n" for c in cues)

cues = [{"text": ">> ALICE: Hi, my name is Alice.", "start": 0.599, "end": 4.16}]
print(to_srt(cues)); print(to_sbv(cues))

Which should you upload, and when should you burn instead?

Upload .srt or .sbv when you want viewers to be able to turn captions off. Pick .vtt or .ttml only if you need the positioning or styling the page says they support, and note that YouTube limits WebVTT styling to bold, italic and underline. Save the file as plain UTF-8, because the page requires it for both basic formats.

A caption file never changes what a viewer sees in a player that ignores tracks. When the look must be the same everywhere, or the clip will be reposted elsewhere, a burned-in caption job takes the same cues. Standalone video captions take cues and burn exactly that copy; a job for a clip up to 60 seconds is $0.20, per the video captions docs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume