SRT or WebVTT? What Vimeo, Spotify, Apple and Cloudflare accept
Vimeo, Spotify and Apple Podcasts take SRT or WebVTT; Cloudflare Stream documents WebVTT. A table read 2026-10-03 and a script writing both from Sume segments.

Vimeo, Spotify and Apple Podcasts each say they accept both SRT and WebVTT files, while Cloudflare Stream's captions page documents WebVTT only (read 2026-10-03). Sume's speech-to-text returns timed segments as JSON, not either file format, so you write the file yourself, and writing WebVTT is the safer default because every platform on this list accepts it.
The table and a script that writes both formats from the same segments follow.
Which platforms accept SRT and which WebVTT?
Each row below comes from the vendor's own help or developer page, read on 2026-10-03. Where a page is silent on a format I say so instead of guessing, since a missing mention is not a rejection.
| Platform | SRT | WebVTT | Other limits on the page |
|---|---|---|---|
| Vimeo | Supported | Supported, recommended | UTF-8 encoding required |
| Spotify | Supported | Supported | Max 5MB; timestamps needed to sync |
| Apple Podcasts | Supported | Supported | Subscriber episodes upload in Apple Podcasts Connect |
| Cloudflare Stream | Not mentioned | Documented (upload and fetch) | One caption track per language |
Why prefer WebVTT when you can?
Vimeo recommends WebVTT outright. Cloudflare's examples upload a .vtt file and return WebVTT when you fetch the track. WebVTT also carries things SRT has no syntax for: the W3C specification defines voice spans such as <v Name> and cue settings for position and size (read 2026-10-03), which Apple's page uses when it says a VTT lets you identify speakers. The cost is strictness. WebVTT needs a WEBVTT header line and a period before the milliseconds where SRT uses a comma, and getting those wrong is the usual cause of Vimeo's invalid caption file error, covered in the Vimeo post.
How do I write both from Sume segments?
Sume's STT 1.0 segments are gapless, with each end equal to the next start, and Vimeo errors when a cue starts at or before the previous one ends, so the script trims one millisecond off each end. It writes UTF-8, formats the separator for each format, and produces both files from one list. Segments come from data.result.segments of a finished transcription (docs cover the video route's data.transcript.segments). I ran the output through ffmpeg in both directions as a syntax check, which is not the same as the platform's own validator, so upload one test file first.
segments = [ # data.result.segments from STT 1.0 (gapless)
{"text": "Welcome back.", "start": 0.0, "end": 1.8},
{"text": "Today: captions.", "start": 1.8, "end": 3.9},
]
def stamp(sec, sep):
ms = round(sec * 1000)
h, ms = divmod(ms, 3_600_000)
m, ms = divmod(ms, 60_000)
s, ms = divmod(ms, 1000)
return f"{h:02}:{m:02}:{s:02}{sep}{ms:03}"
def render(segs, vtt):
sep = "." if vtt else ","
out = ["WEBVTT", ""] if vtt else []
for i, seg in enumerate(segs, 1):
end = seg["end"] - 0.001 # keep cues from touching
out.append(str(i))
out.append(f"{stamp(seg['start'], sep)} --> {stamp(end, sep)}")
out += [seg["text"], ""]
return "\n".join(out)
for ext, vtt in (("srt", False), ("vtt", True)):
with open("captions." + ext, "w", encoding="utf-8") as f:
f.write(render(segments, vtt))
print(render(segments, False))
What else do the platforms check?
Format is only the first gate. Vimeo's troubleshooting page says only one caption track can be active per language and type, and that its own auto captions and your uploaded captions cannot be active together, so an uploaded file you cannot see usually means the wrong track is switched on. Spotify's page says a transcript without timestamps will not sync during playback, and that Spotify has no in-app editor, so a fix means downloading the file, editing it and uploading it again. Apple asks that transcripts be free of spelling and punctuation errors and include all segments of the episode, and attributes the displayed transcript to the show (all read 2026-10-03).
Cloudflare Stream treats each language as unique on a video. A second upload for the same language tag therefore targets the same track rather than adding another. Speech-to-text output is a good first draft, not a finished transcript, so read it against the audio before you upload anywhere that shows the text to listeners.
What can't Sume do here?
Four limits to plan around:
- Sume does not export SRT or WebVTT. STT returns JSON segments, so the conversion above is your code.
- STT 1.0 has no public speaker labels, so a VTT with voice names needs turns you supply yourself.
- Burning goes the other way: the caption job does not accept SRT uploads, only
cuesorsegmentsJSON (video captions). - A single STT call is capped at 600 seconds of audio, so longer recordings must be split and the timestamps offset when you stitch the file.
What does it cost?
Speech-to-text is $0.01 per audio minute on the STT 1.0 endpoint; confirm live pricing in GET /v1/catalog. The file conversion is free, since it runs on your machine. For the other direction, see burning an SRT into video and generating an SRT from a video.
Sources
- Vimeo Help: add captions or subtitles (read 2026-10-03)
- Vimeo Help: troubleshooting caption and subtitle issues (read 2026-10-03)
- Spotify for Creators: managing episode transcripts (read 2026-10-03)
- Apple Podcasts: transcripts (read 2026-10-03)
- Cloudflare Stream: adding captions (read 2026-10-03)
- W3C WebVTT specification (read 2026-10-03)
- Video inspect
- Video captions
Related posts
More in Media tools
- Swap the model in a clothing video with AI: H3 Max Recast
Recast swaps the person in a clip for a person from a photo, 5 to 30 seconds. What it keeps, what the docs leave open about the garment, and the Sume call.
- Turn a ChatGPT try-on image into a video with Seedance 2.5
Saved a try-on image to your ChatGPT Library? Host it, then use it as the first frame of a 9:16 clip on seedance-2.5 through POST /v1/videos. A working script.
- Turn a hum into music with AI: what Sume takes as input
Stability says hum-to-steer is coming. Sume's music API takes text and one optional image, not audio. Here is how to describe a hummed tune in a prompt.
- Vertical video subtitles: BBC's 3-line rule and Sume settings
BBC guidance for 9:16 subtitles: up to 3 lines, 90% width, placed a little high. How to set safe_width_ratio and anchor_ratio on Sume burned-in captions.
Written by Sume