LinkedIn video captions: English SRT only vs burned-in text
LinkedIn's Videos API takes one English caption file, shown with SRT. Sume video-captions burns your cues text into the video and does not accept SRT uploads.

They are two different things. LinkedIn's page says "Each video can include only one caption file, and only English-language captions are supported," while Sume's POST /v1/video-captions burns words into the picture and lists SRT uploads as unsupported, so pass cues instead.
What does LinkedIn's API take?
The page's upload sample uses a caption with "format": "SRT". It limits each video to one caption file, in English.
| Item | LinkedIn caption file | Sume burned-in captions |
|---|---|---|
| Where the text lives | Caption file | Pixels in the MP4 |
| Languages | English only | language hint; speech-to-text or your cues |
| SRT | Shown in the sample | Unsupported as input |
| Cost | Not quoted here | $0.20 per job up to 60 s |
How do I burn my own text?
Pass cues with text, start and end in seconds. That skips speech-to-text and burns exactly that copy. script_text, words, cues and segments are mutually exclusive. design.placement.anchor_ratio sets the line's centre as a fraction of frame height.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: li-caption-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/clean.mp4",
"cues": [{ "text": "Q3 results in 30 seconds", "start": 0, "end": 3 }]
}'Should I do both?
They can coexist: the caption file is what LinkedIn's API takes, and burned-in text is part of the picture. Burned-in words cannot be turned off by the viewer, so keep them short. If LinkedIn needs an English file, write the SRT yourself from the same cues.
What is unclear?
The page does not say what happens with a non-English caption, beyond "only English-language captions are supported." Sume's caption job is a render, so re-run it if the wording changes.
How do cues line up with a caption file?
Your cues are the source of truth for both. Write them once with text, start and end in seconds; burn them with Sume; and convert the same list to a caption file for LinkedIn in your own code. That keeps the two copies from drifting apart. Sume does not write or read SRT, so the conversion is yours.
Sources
Related posts
More in Use cases
- LinkedIn Videos API thumbnail: pick a still from the clip first
LinkedIn's Videos API page says a system thumbnail may be added if you upload none. Sume video-frames returns jpeg stills at times you choose.
- Turn product photos into a silent video with Timeline static holds
Timeline 1.0 treats a still as a static hold, so product photos become a silent video with fades. Shopify lists MP4 up to 10 minutes; here is the request.
- Pull product stills from a video with the video-frames API
Video frames returns jpeg or png stills from a Sume-hosted clip at set seconds. Google lists 500 x 500 minimum, so extract from the clip before captions.
- Shopify's recommended 2048 x 2048 product image with gpt-image-2.5
Shopify.dev recommends 2048 x 2048 px for product images. That is a valid custom image_size on Sume's gpt-image-2.5: 2048 is a multiple of 16 and 4,194,304 px.
Written by Sume