Caption files: YouTube UTF-8, LinkedIn text-only SRT, Sume burns in

YouTube takes SRT, SBV and more as UTF-8; LinkedIn wants text-only SRT. Sume burns captions in and exports no SRT, so build the file from segments.

4 min readSume
All posts

YouTube accepts caption files in several formats including .srt and .sbv, as plain UTF-8, and LinkedIn video ads want captions as an SRT file with text only. Sume does not export an SRT, and its caption model does not take one. It burns text into the picture. For a sidecar file, take the sentence segments from video inspect and write the SRT yourself.

What the platforms ask for

The facts below are from the YouTube caption file page and the LinkedIn video ads page, both read on 2026-10-08.

Caption file rules, read 2026-10-08
PlatformAcceptedRule
YouTubesrt, sbv or sub, mpsub, lrc, capPlain text, UTF-8; basic types carry no style information
LinkedIn video adssrtText only

What Sume gives you

Video captions burns styled text onto a public HTTPS video for $0.20 per job on videos up to 60 seconds. It does not accept SRT upload. Video inspect returns a transcript when you ask for it, and with segmentation set to sentence each segment has index, text, start, end and duration_seconds. Speech-to-text is $0.01 per audio minute plus Modal compute.

That is enough to build the file. The script below turns segments into SRT text. It uses sample data in the same shape; replace it with the segments you read from a finished inspect job.

def stamp(t):
    ms = round(t * 1000)
    h, ms = divmod(ms, 3600000)
    m, ms = divmod(ms, 60000)
    s, ms = divmod(ms, 1000)
    return f"{h:02}:{m:02}:{s:02},{ms:03}"

segments = [
    {"index": 0, "text": "Welcome back.", "start": 0.0, "end": 1.4},
    {"index": 1, "text": "Three specs changed.", "start": 1.4, "end": 3.2},
]
lines = []
for n, seg in enumerate(segments, 1):
    lines.append(str(n))
    lines.append(stamp(seg["start"]) + " --> " + stamp(seg["end"]))
    lines.append(seg["text"])
    lines.append("")
open("captions.srt", "w", encoding="utf-8").write("\n".join(lines))

Which route for which job

Use the sidecar SRT when the platform shows its own caption control, as YouTube does, because viewers can turn it off and the platform can translate it. Use burned-in captions when the placement has no caption control, or when you want a styled look. The two are not exclusive: you can upload the SRT and also ship a burned-in version for placements that need it.

Name the file after the video so the pair is easy to match later. Keep the file plain. YouTube's page says the basic types hold no style information, and LinkedIn asks for text only, so leave out tags and positioning.

  • Sidecar SRT: free to build once you have segments; no Sume caption job.
  • Burned-in: $0.20 per job up to 60 seconds, style such as punch or tiktok-green.
  • Both: inspect transcript once, then reuse the same segments for SRT and for cues.

Check the encoding

Write the file as UTF-8 with no byte order surprises, which Python's encoding argument above handles. Open it in a text editor and confirm that accented or Hangul text reads correctly before you upload, since a wrong encoding is the most common reason a caption file is refused.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume