SRT vs VTT: the difference between the two subtitle formats
SRT and VTT are both plain-text subtitle files. VTT adds a WEBVTT header, dot milliseconds, positioning and styling, and is what HTML5 players read.

SRT and VTT are both plain-text subtitle files that pair each line of text with a start and an end time. SRT (SubRip) is the basic one: numbered blocks with a comma before the milliseconds. VTT (WebVTT) starts with a WEBVTT header, puts a dot before the milliseconds, and adds cue settings for position and optional CSS styling. It is the format the HTML <track> element reads.
So pick the format your destination asks for: VTT for a web player, either one for YouTube. The format rules below come from MDN's WebVTT format and <track> element pages and YouTube Help's supported subtitle files page, read on 2026-09-28.
What is the difference between SRT and VTT?
The SRT column follows the SubRip example and rules on YouTube's help page; the VTT column follows MDN.
| Part | SRT (.srt) | VTT (.vtt) |
|---|---|---|
| First line | The first block's number | WEBVTT, the only required part of the file |
| Cue numbers | Each block starts with a number | Optional identifiers; numbering them is common |
| Time format | 00:00:00,599 --> 00:00:04,160 | mm:ss.ttt or hh:mm:ss.ttt, joined by --> |
| Position | YouTube points to its advanced formats, not SRT, for positioning | Cue settings such as line, position, size, align |
| Styling | YouTube recognizes no style markup | STYLE blocks with the ::cue pseudo-element |
| Encoding | Plain UTF-8 on YouTube | UTF-8; MIME type text/vtt |
HTML5 <track> | Not the format MDN documents | The format <track> files use |
What does an SRT file look like?
A number, a time line, the text, then a blank line before the next block. This follows the SubRip example on YouTube's help page:
1
00:00:00,599 --> 00:00:04,160
Three steps to a captioned video.
2
00:00:04,160 --> 00:00:06,770
First, write the script.What does a VTT file look like?
The same two cues in WebVTT. MDN lists five cue settings (vertical, line, position, size, align); here line:-1 places the first cue at the bottom, since MDN counts negative line numbers from the bottom up:
WEBVTT
1
00:00:00.599 --> 00:00:04.160 line:-1
Three steps to a captioned video.
2
00:00:04.160 --> 00:00:06.770
First, write the script.Which one should I use?
- A website or web app: VTT. MDN says
<track>files are formatted in WebVTT, withkindvalues such assubtitlesandcaptions. - YouTube: either. Its help page lists SubRip (.srt) as a basic format, with only basic files supported and no style info recognized, and WebVTT (.vtt) as an advanced format in initial implementation: positioning is supported, and styling is limited to
<b>,<i>, and<u>. - Another platform or player: use the format its own upload page names.
How do I convert SRT to VTT?
Add WEBVTT and a blank line at the top, and change the comma before the milliseconds to a dot in every time line. The SRT numbers can stay: in WebVTT they become cue identifiers, which MDN says do not have to be unique. Save the result as UTF-8 with a .vtt extension. Going from VTT to SRT, reverse it: remove the header, any STYLE, REGION, or NOTE blocks, and the cue settings, number every block, and change each dot before the milliseconds back to a comma. YouTube recognizes no style markup in SRT files, so styling does not carry over.
Do burned-in captions need an SRT or VTT file?
No. Burned-in captions are drawn into the picture, so there is no file to ship; Hard subtitles vs soft subtitles explains the trade-off. Sume's captions API burns captions but takes no SRT upload: send phrase-level cues with text, start, and end in seconds instead, as How to burn an SRT file into a video shows. It returns the captioned video, not a subtitle file. To write an .srt or .vtt from speech, see How to generate an SRT file from a video.
Sources
Related posts
More in Media tools
- Steam capsule sizes: header, small, main, and vertical
Steam's capsules are 920×430 (header), 462×174 (small), 1232×706 (main) and 748×896 (vertical). What each may show, and how to make the art at size.
- What is text-based video editing? How transcript cuts work
Text-based video editing cuts a video by deleting words from its transcript. How word timings become a cut list, and how to build it with Sume's API.
- Types of cuts in video editing: hard, jump, J, L, and more
The main types of cuts in video editing: hard, jump, J and L, match, cutaway, cross-cut, and smash cut. What each does, and how to build them.
- Vertical video resolution: 9:16 sizes from 720p to 4K
Vertical video resolution is 1080×1920 for 1080p, 720×1280 for 720p, and 2160×3840 for 4K. How to size any 9:16 frame, and what Sume can render.
Written by Sume