Bluesky video captions: WebVTT sidecar vs burned-in
Bluesky video embeds take WebVTT caption files next to the video. Sume burns captions into pixels, so build the VTT yourself from sentence segments.

A Bluesky video post can carry captions as separate text/vtt files next to the video, and Sume does not produce those. Sume's caption endpoint burns text into the pixels. To get a sidecar track, transcribe the clip with video_inspect and write the WebVTT yourself from the sentence segments.
Bluesky facts are from the app.bsky.embed.video lexicon and Sume facts from Video captions and Video inspect, read 2026-09-30.
What does the Bluesky video embed accept for captions?
The lexicon lists a captions array with at most 20 entries. Each entry has a lang and a file, and the file must be text/vtt with a maxSize of 20000 bytes. The video itself is an mp4 blob of up to 300 MB.
| Field | Rule in the lexicon |
|---|---|
video | video/mp4, maxSize 300000000 |
captions | Array, maxLength 20 |
captions[].lang | Language code, required |
captions[].file | text/vtt, maxSize 20000 |
alt | Alt text string |
Does Sume produce a VTT or SRT file?
No. POST /v1/video-captions takes a public HTTPS video_url and returns a captioned video. The docs state that SRT uploads are unsupported, and you pass phrase text as cues with text, start and end instead. The result is a video with the words drawn on it, so it is a different thing from a sidecar track that a viewer can turn off or switch by language. For the file-format difference see SRT vs VTT.
How do I build the VTT from Sume segments?
Call POST /v1/video-inspect with transcribe: true and segmentation: { mode: "sentence" }. The result carries sentence segments[] with text, start and end in seconds. Keep each file under 20000 bytes for Bluesky, and use one file per language.
const pad = (n, w = 2) => String(n).padStart(w, "0");
const ts = (s) =>
pad(Math.floor(s / 3600)) + ":" +
pad(Math.floor((s % 3600) / 60)) + ":" +
(s % 60).toFixed(3).padStart(6, "0");
export function segmentsToVtt(segments) {
const cues = segments.map(
(c) => ts(c.start) + " --> " + ts(c.end) + "\n" + c.text.trim(),
);
return "WEBVTT\n\n" + cues.join("\n\n") + "\n";
}
// const vtt = segmentsToVtt(inspect.transcript.segments);
// if (Buffer.byteLength(vtt) > 20000) throw new Error("too big for Bluesky");Should I burn captions or ship a sidecar?
Use the sidecar when you want viewers to toggle captions and want several languages on one video. Burn captions when the clip will travel to places that ignore sidecar files. You can do both: post the original video with VTT files on Bluesky, and make a burned-in copy with cues for other destinations. Transcription is billed at the STT rate described in the Video inspect docs; check GET /v1/catalog for live pricing.
Sources
Related posts
More in Use cases
- Break down a UGC ad by scene: $0.30 analysis or free stills?
To study a UGC-style ad, video inspect gives free stills and a $0.01 per minute transcript; the legacy analysis adds typed scenes for $0.30. What each returns.
- Virtual staging disclosure: California AB 723 and originals
California AB 723 asks listings with digitally altered photos to say so and link to the unaltered originals. How to keep original and edit as two URLs.
- California AI Transparency Act date, and latent vs manifest marks
AB 853 delays the California AI Transparency Act to August 2, 2026. What latent and manifest mean in the bill, and what Sume's docs say it ships.
- Cut one product video into feed cutdowns with the video-trim API
Video trim cuts a [start, end) range from a Sume-hosted clip into a new MP4 at $0.02 per job. Meta's Facebook Feed page lists 1 second to 241 minutes of length.
Written by Sume