Vimeo "Invalid Caption File": build clean WebVTT from Sume
Vimeo rejects a cue that starts at the previous cue's end. Sume sentence segments touch, so trim 1 ms, write UTF-8 WebVTT and upload. Script included.

Vimeo's "Invalid Caption File" error on a WebVTT upload usually means one of a short list of syntax problems, and the one that bites files built from a transcript is cue timing: Vimeo rejects a cue whose start is equal to or earlier than the previous cue's end (read 2026-10-03). Sume's sentence segments are gapless by design, so each segment ends exactly where the next begins, and a naive conversion trips that rule.
The fix is to pull every cue end back by one millisecond, write the file as UTF-8 with a WEBVTT header, and upload it as a WebVTT file. Below is the full list of what Vimeo checks, then a script that builds a clean file from a Sume transcript.
What does Vimeo say makes a caption file invalid?
Vimeo supports SRT and WebVTT for captions and subtitles, recommends WebVTT, and requires UTF-8 encoding so special characters display correctly (read 2026-10-03). Its troubleshooting page separates two errors. "Unexpected Text Track Upload Type" means the file type is not SRT or WebVTT, for example a .txt or .csv. "Invalid Caption File" means the type is right but the contents have formatting problems such as HTML code or improper syntax.
For the second error Vimeo's steps are to add a WEBVTT header followed by a line break, replace commas in timestamps with periods if the file started as SRT, and run the text through a WebVTT validator, then fix whatever it reports.
| Vimeo rule | What goes wrong | How to meet it from Sume output |
|---|---|---|
| Header line WEBVTT | SRT-derived files have no header | Write WEBVTT and a blank line first |
| Period before milliseconds | SRT uses a comma | Format as HH:MM:SS.mmm |
| Both timestamps on one line | A line break after the start time | Emit start --> end on a single line |
| Start must come after the previous end | Equal times are rejected | Shorten each end by 1 ms; Sume segments touch |
| UTF-8 | Other encodings garble characters | Open the file with encoding utf-8 |
Why do Sume segments trip the start-after-end rule?
When you request segmentation.mode: "sentence" on a transcript, Sume returns gapless sentence segments: each segments[i].end equals segments[i+1].start, with a default 70 ms boundary lead carried past the last word of a sentence (video inspect, OpenAPI). That is exactly what you want for burned-in caption lines, and exactly what Vimeo's example of a failing file shows: the end of one cue equal to the start of the next, at 00:18.000.
The WebVTT standard itself does not forbid touching cues, so the same file may play fine elsewhere and still fail Vimeo's upload. Treat the one-millisecond trim as a Vimeo-specific normalisation and apply it only to the file you send there.
How do I build the file from a Sume transcript?
Vimeo wants a video you already host, so the video here is a clip in your Sume workspace. Import anything from elsewhere first with POST /v1/media-imports; the inspect route does not fetch open-internet URLs (media inputs). Transcription is billed at $0.01 per audio minute under the current public rate, and duration_seconds is capped at 600 as a reservation hint, so cut longer videos into parts first with video trim.
The script submits an inspect with frames: false so no stills are made, polls until it finishes, and writes talk.vtt.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def stamp(s):
ms = round(s * 1000)
return f"{ms // 3600000:02d}:{ms // 60000 % 60:02d}:{ms // 1000 % 60:02d}.{ms % 1000:03d}"
def to_vtt(segs):
out = ["WEBVTT", ""]
for i, g in enumerate(segs):
end = g["end"] if i + 1 == len(segs) else min(g["end"], segs[i + 1]["start"] - 0.001)
out += [f"{stamp(g['start'])} --> {stamp(end)}", g["text"].strip(), ""]
return "\n".join(out)
body = {"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "frames": False,
"transcribe": True, "duration_seconds": 120, "mode": "async",
"segmentation": {"mode": "sentence"}}
job = requests.post(f"{API}/v1/video-inspect", json=body,
headers={**H, "Idempotency-Key": "vimeo-vtt-001"}).json()["data"]
while True:
vi = requests.get(f"{API}/v1/video-inspect/{job['video_inspect_id']}", headers=H).json()["data"]
if vi["status"] in ("completed", "failed", "canceled"):
break
time.sleep(2)
if vi["status"] != "completed":
raise SystemExit(f"inspect ended as {vi['status']}")
open("talk.vtt", "w", encoding="utf-8").write(to_vtt(vi["transcript"]["segments"]))
What should I check before uploading?
A two-minute pass catches most rejections before Vimeo does, and it keeps you from re-uploading a file that will fail the same way again. Run through this list every time you regenerate the file, because a changed transcript changes the cue boundaries.
- Open the file in a text editor: the first line must read WEBVTT with nothing before it.
- Paste it into a WebVTT validator, as Vimeo suggests, and read the first reported line.
- Upload, then pick the language and the Type (Subtitle or Caption) from the dropdowns, and toggle the track on.
- Remember Vimeo allows only one active text track per language and type, and your own file cannot be active together with Vimeo's autogenerated captions.
- Read the transcript once: Sume returns speech-to-text wording, and a human pass on names and jargon is still worth it.
When is burning captions in the better choice?
A sidecar file keeps captions switchable and searchable, which is what Vimeo's player expects. It carries no styling from Sume, though. If the clip is going to social feeds as well, burn captions with `POST /v1/video-captions` and keep the VTT for the Vimeo page. The caption endpoint takes no SRT or VTT upload; to burn existing wording, pass its text and times as cues, and the job skips transcription.
The two are separate jobs with separate prices: a caption job is $0.20 per job for videos up to 60 seconds under the current fixed estimate, so confirm live pricing in the catalog before batching.
Sources
Related posts
More in Media tools
- WebVTT cue settings (line, size) vs Sume anchor_ratio and width
WebVTT line:78%,center and size:90% have close matches in Sume's caption design fields. A converter script, and what per-cue settings a burned render loses.
- Where to break subtitle lines: BBC rules and Sume caption cues
The BBC says one sentence per subtitle and no article-noun splits. Sume sentence segments cover the first rule; cues with a line break cover the rest.
- Who is speaking in subtitles: BBC colours, dashes, and Sume cues
WCAG 1.2.2 wants speaker identification in captions. The BBC prefers colour, then dashes or labels. What Sume's burned-in captions can do and a dash script.
- YouTube Auto-sync captions vs Sume script_text: track or burned in
YouTube Auto-sync times your transcript into a caption track. Sume script_text aligns your wording into burned-in captions. What each needs and which to use.
Written by Sume