X API subtitles media_category: SRT sidecar vs burned-in captions
X lists `subtitles` as a media_category for sidecar files. Sume burns captions into the picture and does not take SRT uploads. Pick the path that fits.

X's media upload page lists subtitles as one of its media categories, which is the category for a separate subtitle file. Sume does the other thing: POST /v1/video-captions burns the words into the video frames, and its docs say SRT uploads are unsupported.
X facts are from its media best-practices page, and Sume facts from the Video captions and Video inspect docs, all read 2026-10-01.
What does X say about the subtitles category?
The page defines the media category parameter as the use case of the file, passed in the INIT request of the upload flow. It lists tweet_image, tweet_video, tweet_gif, amplify_video, dm_image, dm_video, dm_gif and subtitles. It also warns that using the wrong category is a common reason an upload succeeds and Post create then fails. The page does not describe the subtitle file format, so check X's subtitle-specific docs for that before you build.
What does Sume produce instead?
Sume burns captions onto an existing public video URL and returns a job-backed captioned video. The docs say SRT uploads and provider task ids are unsupported; for phrase-level text you pass cues or segments with text, start and end instead. That burned-in file goes up as an ordinary tweet_video.
| Question | X sidecar `subtitles` | Sume video-captions |
|---|---|---|
| Where do the words live? | A separate file | Burned into the picture |
| Can a viewer switch them off? | Not stated on the X page | No, they are pixels |
| SRT input | Not described on the page | Unsupported; use cues or segments |
| Upload category | subtitles | tweet_video for the finished clip |
Can I still get a transcript for a sidecar file?
Yes, as raw material. POST /v1/video-inspect with transcribe: true runs Sume STT 1.0 and returns text and words[]; segmentation.mode: "sentence" also returns sentence segments[] that are shaped like caption lines. Writing the file X expects from those times is your code, since Sume does not emit SRT. For a related comparison on another platform, see LinkedIn SRT versus burned-in captions.
How do I burn captions with cues?
Send the copy and times you already have. Authored cues skip speech-to-text.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: x-captions-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"cues": [{ "text": "Launch day", "start": 0, "end": 2.5 }]
}'Which one should I choose for X?
Choose burned-in when the post must read with sound off and you want one file to upload. Choose a sidecar when viewers need to toggle or translate the text, and build it from the inspect transcript. Note that the captions route needs a fetchable public HTTPS video_url.
Sources
Related posts
More in Developers
- X DM video limit: 140 seconds and 512 MB, and how to trim to it
X lists DM video (dm_video) at 140 seconds and 512 MB by default. Cut a clip to length with Sume video-trim, which takes 0.2 to 900 seconds.
- X video minimum 0.5 seconds: keep Sume trim cuts at 0.5 s or more
X needs Post video of at least 0.5 seconds; Sume video-trim allows cuts down to 0.2 seconds. Set duration to 0.5 or more for X, with the upper caps.
- X video audio must be AAC-LC, not HE-AAC: check an AI video
X requires AAC Low Complexity audio, mono or stereo, with H264 High Profile and YUV 4:2:0. What Sume's exact trim keeps and re-encodes, and what it leaves.
- X video max 1280x1024, ratio 1:3 to 3:1, 60 fps: conform a clip
X video must be 32x32 to 1280x1024, 60 fps or less, ratio 1:3 to 3:1. Read literally, a default 1080x1920 Timeline file is over; conform with video-trim.
Written by Sume