LinkedIn wants a text-only SRT; Sume burns captions, so build one

LinkedIn video ad captions must be a text-only SRT. Sume burns captions into pixels and has no SRT export. Build the file from video-inspect segments.

4 min readSume
All posts

If LinkedIn asks for captions as a separate file, Sume's burned-in captions are not that file. LinkedIn's help page says captions "must be in SRT format" and that you should "only include text in your SRT file". Sume's caption tool renders text into the video frames and the docs I read describe no SRT output or upload. To get an SRT, ask video_inspect for a transcript with sentence segments and write the file yourself.

Two ways to caption a LinkedIn video

Pick by whether you need a toggleable track or a permanent look.

Burned captions vs SRT for LinkedIn (vendor page read 2026-10-08)
RouteWhat you getCost per the Sume docs
video_captions burn-inText drawn onto the frames; cannot be turned off by the viewer$0.20 per job, priced for videos up to 60 seconds
video_inspect transcriptText, words and optional sentence segments you turn into an SRT$0.01 per audio minute, plus the inspect compute
LinkedIn SRT uploadA text-only .srt file attached to the adSet in LinkedIn, not Sume

Steps

Import the clip, then call POST /v1/video-inspect with transcribe: true, a language_code, and segmentation: {"mode": "sentence"}. The docs say sentence segments come back with no gaps, in the shape of caption lines. Open one real response to confirm the field names before you map them, because I did not verify them.

Once you have start, end and text for each line, a small formatter makes the file. This one runs as is.

def ts(sec):
    ms = round(sec * 1000)
    h, ms = divmod(ms, 3600000)
    m, ms = divmod(ms, 60000)
    s, ms = divmod(ms, 1000)
    return f"{h:02}:{m:02}:{s:02},{ms:03}"

def to_srt(lines):
    out = []
    for i, (start, end, text) in enumerate(lines, 1):
        out.append(f"{i}\n{ts(start)} --> {ts(end)}\n{text}\n")
    return "\n".join(out)

sample = [(0.0, 2.4, "Holiday sale starts today."),
          (2.4, 5.1, "Free shipping on every gift.")]
print(to_srt(sample))

Save the output as a .srt file and keep it text only, as LinkedIn asks. If the transcript is wrong, fix the text before you write the file.

Which route to choose

Use the SRT route when LinkedIn is the only place the video will run and you want captions viewers can switch off. The cost is the transcript at $0.01 per audio minute plus the compute, and your own time to review it. Use burned captions when the same file goes to several places that do not take caption files, because the text travels inside the picture.

You can do both. Burn a short headline or offer line as an overlay with cues, and upload an SRT for the spoken words. Then the two do not duplicate each other. Keep the SRT text only, as LinkedIn's page asks, with no styling tags.

Always read the transcript before you publish. The docs describe speech-to-text, not a human review, and brand names and product terms are where it is most likely to slip.

Budget and detail

A practical budget for ten 30-second LinkedIn videos looks like this. Ten transcript requests cost about $0.01 per audio minute each, so roughly five cents of speech-to-text at half a minute per clip, plus the inspect compute, which the docs bill by container seconds and which I could not price from them. Ten burned-in caption jobs would be $2.00 at $0.20 each. The SRT route is therefore the cheaper one when you only need the file, and the burn route is the one that guarantees the text is visible.

What Sume does not do

Sume does not generate or accept an SRT for you, and it does not check a file against LinkedIn's caption processing. A silent clip returns inspect_source_has_no_audio, so check probe.has_audio first.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume