LinkedIn video captions: English SRT only vs burned-in text

LinkedIn's Videos API takes one English caption file, shown with SRT. Sume video-captions burns your cues text into the video and does not accept SRT uploads.

4 min readSume
All posts

They are two different things. LinkedIn's page says "Each video can include only one caption file, and only English-language captions are supported," while Sume's POST /v1/video-captions burns words into the picture and lists SRT uploads as unsupported, so pass cues instead.

What does LinkedIn's API take?

The page's upload sample uses a caption with "format": "SRT". It limits each video to one caption file, in English.

Caption paths (read 2026-09-30). Sume rows: https://docs.sume.com/models/video-captions
ItemLinkedIn caption fileSume burned-in captions
Where the text livesCaption filePixels in the MP4
LanguagesEnglish onlylanguage hint; speech-to-text or your cues
SRTShown in the sampleUnsupported as input
CostNot quoted here$0.20 per job up to 60 s

How do I burn my own text?

Pass cues with text, start and end in seconds. That skips speech-to-text and burns exactly that copy. script_text, words, cues and segments are mutually exclusive. design.placement.anchor_ratio sets the line's centre as a fraction of frame height.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: li-caption-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/clean.mp4",
    "cues": [{ "text": "Q3 results in 30 seconds", "start": 0, "end": 3 }]
  }'

Should I do both?

They can coexist: the caption file is what LinkedIn's API takes, and burned-in text is part of the picture. Burned-in words cannot be turned off by the viewer, so keep them short. If LinkedIn needs an English file, write the SRT yourself from the same cues.

What is unclear?

The page does not say what happens with a non-English caption, beyond "only English-language captions are supported." Sume's caption job is a render, so re-run it if the wording changes.

How do cues line up with a caption file?

Your cues are the source of truth for both. Write them once with text, start and end in seconds; burn them with Sume; and convert the same list to a caption file for LinkedIn in your own code. That keeps the two copies from drifting apart. Sume does not write or read SRT, so the conversion is yours.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume