Dialogue-only subtitles vs CC: Netflix's two English tracks, burned
Netflix offers English (dialogue only) and English (CC) as separate options. How the two differ, and how to burn either version of one clip with Sume cues.

Dialogue-only subtitles show what people say and nothing else. Closed captions for the deaf and hard of hearing (SDH/CC) add speaker names and sound cues such as [door slams]. Netflix now lists both as separate English choices, so a viewer picks the density they want. If you burn text into a video you have to pick for them, and with Sume you can build both versions from one script and render each as its own job.
The facts about Netflix's two tracks below come from its announcement of the new subtitle option, read 2026-10-03.
What are the two English options on Netflix?
Netflix says the language picker on a new title now shows two entries: English, which shows only the spoken dialogue, and English (CC), which includes dialogue plus audio cues like [door slams]. Before this, a viewer who wanted original-language subtitles had to turn on SDH/CC, which also carried cues like [phone buzzing] or [dramatic music swells] and speaker names.
Netflix adds that the dialogue-only option will be available on all new Netflix originals in every language it offers, in addition to SDH/CC, and that viewers can change the size and font of their subtitles. It also reports that nearly half of all viewing hours in the US happen with subtitles or captions on.
How do the two versions differ?
The split is a content decision, not a styling one. The table lists what each version carries, per Netflix's description, and what that means when you author the text yourself.
| Version | What Netflix says it shows | What you author for a burned-in version |
|---|---|---|
| English | Only the spoken dialogue | One cue per spoken phrase, no brackets, no speaker labels |
| English (CC) | Dialogue plus audio cues like [door slams] | The same cues plus bracketed sound cues |
| SDH/CC before the split | Also speaker names and cues like [phone buzzing] | Add a speaker label where the speaker is not obvious |
Which one should a burned-in caption be?
Burned-in text cannot be toggled, so it should match the audience that cannot opt out. For a muted autoplay feed, dialogue only keeps the picture clear. For a clip that deaf or hard-of-hearing viewers must follow, the sound cues carry meaning, and the version with cues is the one to burn. If you serve both audiences, render two files and publish each where it fits, instead of one compromise.
Sound cues should be short and only for sounds that matter. A doorbell that drives the scene earns a cue; footsteps you can see do not.
How do I render both versions from one script?
Standalone video captions accept cues, each with text, start and end in seconds, which skip speech-to-text and burn exactly that copy at those times. Keep one list with a kind field, filter it twice, and submit two jobs. Each accepted job for a clip up to 60 seconds is $0.20 of Sume usage, per the video captions docs. script_text, words, cues and segments are mutually exclusive, so send only cues.
import json
script = [
{"kind": "speech", "text": "Did you hear that?", "start": 0.4, "end": 1.8},
{"kind": "sound", "text": "[door slams]", "start": 2.0, "end": 3.0},
{"kind": "speech", "text": "Someone is here.", "start": 3.2, "end": 4.6},
]
def body(kinds, video_url):
cues = [{k: c[k] for k in ("text", "start", "end")}
for c in script if c["kind"] in kinds]
return {"video_url": video_url, "cues": cues}
url = "https://media.sume.com/artifacts/example/clean.mp4"
print(json.dumps(body({"speech"}, url), indent=1))
print(json.dumps(body({"speech", "sound"}, url), indent=1))POST each body to /v1/video-captions with its own Idempotency-Key, then poll the job as described in Jobs and results. A silent clip works with cues, since no speech-to-text runs. If a viewer will see the text on a phone, check the line length first; subtitle limit per line covers max_chars.
Sources
Related posts
More in Use cases
- Diwali 2026 five-day story series: one clip per day, Nov 6 to 10
Diwali 2026 runs five days from Dhanteras on 6 November to Bhai Dooj on 10 November. Make five 6-second stories from one music bed for about $1.88 on Sume.
- Do Microsoft Advertising video ads run in Copilot? Eligible ad types
Ads in Copilot are built from your text and image assets; the eligible list names multimedia, product and logo search ads, not video. Sume makes the 15 s clip.
- Does a video with only background music need captions? W3C's table
W3C says captions are needed when audio carries information: Level A for prerecorded video, AA for live. What that means before you burn captions.
- Does a voiceover make a reused YouTube Short original?
No: YouTube's Oct 1, 2026 update says narration describing what's on screen doesn't make a re-upload original. What to change instead, and how Sume helps.
Written by Sume