Who is speaking in subtitles: BBC colours, dashes, and Sume cues
WCAG 1.2.2 wants speaker identification in captions. The BBC prefers colour, then dashes or labels. What Sume's burned-in captions can do and a dash script.

WCAG 1.2.2 says captions should carry more than dialogue, including speaker identification, and the BBC's subtitle guidelines give a concrete method: give each speaker a colour from white, yellow, cyan or green on a black background, and fall back to dashes or labels when colour is not available (read 2026-10-03). Sume's burned-in captions apply one colour per job and Sume's speech-to-text returns no speaker labels, so on Sume you mark speakers yourself in authored cues, with a dash or a name.
This post lays out the options and a short script that adds BBC-style dashes.
What do WCAG and the BBC require?
The W3C's Understanding page for Success Criterion 1.2.2 defines captions as conveying the spoken dialogue plus non-dialogue audio information, listing sound effects, music, laughter, speaker identification and location. It also notes that open captions, which cannot be turned off, are a form of captions (read 2026-10-03).
The BBC guidelines (v1.2.5) say colour is the preferred way to tell speakers apart, in priority order white, yellow, cyan and green, all on a black background, and a speaker should keep the same colour once assigned. By convention the narrator is yellow. Dashes are described as a legacy technique to be used with colour only when unavoidable, and non-human speakers get labels such as ROBOT: Hello, sir.
| Method | Source | On Sume |
|---|---|---|
| Speaker colour (white, yellow, cyan, green) | BBC 8.3, preferred | One text colour per caption job via design.colors; not per cue |
| Dash at the start of a new speaker's line | BBC 7.3, legacy | Put it in the cue text |
| Label such as ROBOT: | BBC 7.8 | Put it in the cue text |
| Sound and laughter cues | WCAG 1.2.2 | Put [laughs] or similar in the cue text |
| Automatic speaker labels | Not required by either | Not exposed; STT 1.0 returns no speaker labels |
Why can't Sume colour each speaker?
The caption job takes a design.colors block with base, active, stroke, accent, accent_deep and card, and those apply to the whole render (video captions). There is no per-cue colour field in the documented request, and design is unsupported on the punch and tiktok-green styles. The speech-to-text endpoint also has no public speaker labels, because diarization settings are fixed on the server; see the diarization comparison.
So the honest options are text-based. If you need true per-speaker colours, deliver a sidecar track instead of burning. The W3C WebVTT specification's sample interview marks each speaker with a voice span such as <v Roger Bingham> at the start of each cue (read 2026-10-03), which names the speaker in the file and lets a player decide how to show it. The sidecar post shows building one from a Sume transcript.
How do I mark speakers in burned-in cues?
Supply cues with text, start and end in seconds. Cues skip speech-to-text, so you provide the speaker turns, for example from your own notes or an editor's transcript. The script below adds a dash when the speaker changes and keeps a sound cue in brackets. It follows the BBC's dash style, which the BBC calls legacy; with a single colour on screen it is the clearest option left.
import json
turns = [ # (speaker, text, start, end) from your own diarized notes
("host", "Did you see the launch?", 0.0, 1.6),
("guest", "I did. It went smoothly.", 1.6, 3.4),
("guest", "Mostly.", 3.4, 4.0),
("host", "[laughs] Mostly?", 4.0, 5.2),
]
cues, prev = [], None
for who, text, start, end in turns:
# BBC style: a dash starts each new speaker's line
prefix = "- " if who != prev and prev is not None else ""
cues.append({"text": prefix + text, "start": start, "end": end})
prev = who
body = {
"video_url": "https://media.sume.com/artifacts/example/interview.mp4",
"style": "slam",
"cues": cues,
"design": {"colors": {"base": "#FFFFFF"}},
}
print(json.dumps(body, indent=2))
When is a dash enough, and when is it not?
A dash works for short two-person exchanges, where each new line is clearly a reply. It gets ambiguous in a long conversation with three or more voices, because a dash says only that the speaker changed, not who it is. There, prefix the first line of each turn with a name or label, as the BBC does for non-human speakers, and keep the label short so the card still fits the safe width.
A narrator or voice-over is the common case where colour is what viewers expect. Since a Sume job has one colour, one practical split is two renders of the same footage: dialogue clips in the default colour, and narration-only clips in a yellow set through design.colors.base. That keeps the BBC's narrator convention without per-cue colour, at the cost of one job per colour.
What should I check afterwards?
Run one clip and read it on a phone before you batch the rest.
- Each dash starts a new speaker, and a continuing speaker's lines do not carry one.
- Sound cues use brackets and describe the sound, not just the word laughter.
- The text colour has enough contrast with the footage; the BBC puts its text on a black background.
- Names or labels are consistent from the first card to the last.
- The speaker turns come from a person or a diarizing tool, since Sume's STT will not tell you who spoke.
What does it cost?
A caption job is $0.20 per job for videos up to 60 seconds under the current fixed estimate, and cues avoid a second transcription; confirm live pricing in GET /v1/catalog. For sound cues in more depth, see the sound-cue checklist.
Sources
Related posts
More in Media tools
- YouTube Auto-sync captions vs Sume script_text: track or burned in
YouTube Auto-sync times your transcript into a caption track. Sume script_text aligns your wording into burned-in captions. What each needs and which to use.
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
- How to extract frames from a video with the Sume API
POST /v1/video-frames returns stills at the times you name from one Sume-hosted clip, as durable images at source size. The call is unbilled.
Written by Sume