Open vs closed captions vs subtitles: what Sume's captions are
Sume's video captions are open captions burned into the pixels. The docs list no SRT or VTT export, and SRT uploads are unsupported. When that is right.

Sume's POST /v1/video-captions produces open captions: the words are drawn into the video frames, so every viewer sees them and nobody can turn them off. It is not a closed caption track or a subtitle file. The docs describe burning only, and state that SRT uploads and provider task ids are unsupported, so if a platform wants a sidecar file you will need to produce that elsewhere.
The terms get mixed up constantly, and the mix-up decides which tool you pick. Here is a plain split, and a short rule for choosing.
Three things people call captions
Sume's burn-in options cover look and wording: styles such as slam, punch, tiktok-green, korean-ad and black-outline, design overrides for colours, typography, placement, phrasing and motion, and script_text alignment so the burned words match an approved script.
| Kind | Where the text lives | Viewer can switch off | Sume |
|---|---|---|---|
| Open captions | Burned into the pixels | No | Yes: POST /v1/video-captions, $0.20 per job up to 60 s |
| Closed captions | A separate track in the video container or player | Yes | Not described in the docs I read |
| Subtitles | A separate text file (SRT, VTT) or track, usually translation | Yes | Not described; SRT upload unsupported |
| Overlay copy | Authored text at set times, not speech | No | Yes: cues or segments with text, start, end |
When burn-in is the right answer
Choose open captions when the clip will travel without its metadata: reposted, downloaded, embedded in a message, or played muted in a feed. Short vertical video is the common case, because feeds autoplay without sound and a separate track may not be shown by default. The style is also part of the creative, so a bold word-by-word caption is a design choice that a platform's own caption renderer cannot give you.
Choose a sidecar file when the platform requires it, such as an ad system that checks for a caption file, or when viewers need to change language or size. Accessibility rules in some settings call for closed captions specifically. I am not giving legal advice, so check the rule that applies to your case rather than assuming burn-in satisfies it.
A decision helper
The helper below encodes the rule of thumb: ask where the video will play and whether viewers need control. It is deliberately short so you can extend it with your own platforms.
def caption_plan(plays_muted, platform_requires_file, viewer_needs_control):
plan = []
if platform_requires_file or viewer_needs_control:
plan.append("closed/subtitle file (make it outside Sume's burn-in)")
if plays_muted:
plan.append("open captions (Sume POST /v1/video-captions)")
return plan or ["no captions needed? check accessibility rules first"]
cases = {
"muted feed short": (True, False, False),
"ad needing SRT": (False, True, False),
"web course": (False, False, True),
"feed video also hosted": (True, False, True),
}
for name, args in cases.items():
print(name, "->", caption_plan(*args))
Checking a platform's requirement
Before you choose, read the platform's own help page for captions. Some accept a sidecar file and display it as an option; some ignore sidecar files for short vertical video; some require one for a particular ad format. The answer changes over time, so write down the date you read it, as in the tables here.
If you are unsure, burn captions into one version and keep a clean, caption-free master. That way you can add a sidecar track later without re-editing, and you can restyle the burned version through source_caption_id without running speech-to-text again.
Doing both without double work
If you need open and closed captions for the same video, get the wording right once. Write or approve the transcript, pass it as script_text so the burned text matches it, and use the same transcript to produce the sidecar file with whatever tool the destination expects. Then the two versions cannot drift apart.
To restyle without paying for transcription twice, send source_caption_id and the new style. Sume reuses the source video and its word timings, and a restyle is still billed as a render. Microsoft's new transcription models list timestamps and diarisation (read 2026-10-03), which are the inputs a sidecar file builder wants, but check each vendor's output format before you plan around it.
Sources
Related posts
More in Media tools
- OpenShot 4.0.1 Razor tool vs a hosted trim: exact or keyframe
OpenShot 4.0.1 upgrades its Razor cutting tool. Sume's video-trim cuts a clip by start and duration, exact (re-encode) or keyframe (stream copy), at $0.02.
- Performance Max AI video: sourced, enhanced or auto-generated?
Performance Max has three AI video features: sourced (opt-in), enhancements (default on), auto-generated (when you upload none). Which one touches your clip.
- Pick a listing main image from a product clip with video-frames
Extract up to 24 stills from a product clip with POST /v1/video-frames on Sume, choose the sharpest one for a listing, and avoid the media.sume.com trap.
- Picsart upscale API 2x to 8x vs Sume image upscale 1x to 4x
Picsart upscale takes factors 2, 4, 6 and 8 up to 4800x4800. Sume image upscale takes a number from 1 to 4 and returns png, jpg or webp.
Written by Sume