Open vs closed captions vs subtitles: what Sume's captions are

Sume's video captions are open captions burned into the pixels. The docs list no SRT or VTT export, and SRT uploads are unsupported. When that is right.

6 min readSume
All posts

Sume's POST /v1/video-captions produces open captions: the words are drawn into the video frames, so every viewer sees them and nobody can turn them off. It is not a closed caption track or a subtitle file. The docs describe burning only, and state that SRT uploads and provider task ids are unsupported, so if a platform wants a sidecar file you will need to produce that elsewhere.

The terms get mixed up constantly, and the mix-up decides which tool you pick. Here is a plain split, and a short rule for choosing.

Three things people call captions

Sume's burn-in options cover look and wording: styles such as slam, punch, tiktok-green, korean-ad and black-outline, design overrides for colours, typography, placement, phrasing and motion, and script_text alignment so the burned words match an approved script.

Definitions as commonly used; the Sume column is from the Video captions docs, read 2026-10-03.
KindWhere the text livesViewer can switch offSume
Open captionsBurned into the pixelsNoYes: POST /v1/video-captions, $0.20 per job up to 60 s
Closed captionsA separate track in the video container or playerYesNot described in the docs I read
SubtitlesA separate text file (SRT, VTT) or track, usually translationYesNot described; SRT upload unsupported
Overlay copyAuthored text at set times, not speechNoYes: cues or segments with text, start, end

When burn-in is the right answer

Choose open captions when the clip will travel without its metadata: reposted, downloaded, embedded in a message, or played muted in a feed. Short vertical video is the common case, because feeds autoplay without sound and a separate track may not be shown by default. The style is also part of the creative, so a bold word-by-word caption is a design choice that a platform's own caption renderer cannot give you.

Choose a sidecar file when the platform requires it, such as an ad system that checks for a caption file, or when viewers need to change language or size. Accessibility rules in some settings call for closed captions specifically. I am not giving legal advice, so check the rule that applies to your case rather than assuming burn-in satisfies it.

A decision helper

The helper below encodes the rule of thumb: ask where the video will play and whether viewers need control. It is deliberately short so you can extend it with your own platforms.

def caption_plan(plays_muted, platform_requires_file, viewer_needs_control):
    plan = []
    if platform_requires_file or viewer_needs_control:
        plan.append("closed/subtitle file (make it outside Sume's burn-in)")
    if plays_muted:
        plan.append("open captions (Sume POST /v1/video-captions)")
    return plan or ["no captions needed? check accessibility rules first"]

cases = {
    "muted feed short": (True, False, False),
    "ad needing SRT": (False, True, False),
    "web course": (False, False, True),
    "feed video also hosted": (True, False, True),
}
for name, args in cases.items():
    print(name, "->", caption_plan(*args))

Checking a platform's requirement

Before you choose, read the platform's own help page for captions. Some accept a sidecar file and display it as an option; some ignore sidecar files for short vertical video; some require one for a particular ad format. The answer changes over time, so write down the date you read it, as in the tables here.

If you are unsure, burn captions into one version and keep a clean, caption-free master. That way you can add a sidecar track later without re-editing, and you can restyle the burned version through source_caption_id without running speech-to-text again.

Doing both without double work

If you need open and closed captions for the same video, get the wording right once. Write or approve the transcript, pass it as script_text so the burned text matches it, and use the same transcript to produce the sidecar file with whatever tool the destination expects. Then the two versions cannot drift apart.

To restyle without paying for transcription twice, send source_caption_id and the new style. Sume reuses the source video and its word timings, and a restyle is still billed as a render. Microsoft's new transcription models list timestamps and diarisation (read 2026-10-03), which are the inputs a sidecar file builder wants, but check each vendor's output format before you plan around it.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume