Who is speaking in subtitles: BBC colours, dashes, and Sume cues

WCAG 1.2.2 wants speaker identification in captions. The BBC prefers colour, then dashes or labels. What Sume's burned-in captions can do and a dash script.

5 min readSume
All posts

WCAG 1.2.2 says captions should carry more than dialogue, including speaker identification, and the BBC's subtitle guidelines give a concrete method: give each speaker a colour from white, yellow, cyan or green on a black background, and fall back to dashes or labels when colour is not available (read 2026-10-03). Sume's burned-in captions apply one colour per job and Sume's speech-to-text returns no speaker labels, so on Sume you mark speakers yourself in authored cues, with a dash or a name.

This post lays out the options and a short script that adds BBC-style dashes.

What do WCAG and the BBC require?

The W3C's Understanding page for Success Criterion 1.2.2 defines captions as conveying the spoken dialogue plus non-dialogue audio information, listing sound effects, music, laughter, speaker identification and location. It also notes that open captions, which cannot be turned off, are a form of captions (read 2026-10-03).

The BBC guidelines (v1.2.5) say colour is the preferred way to tell speakers apart, in priority order white, yellow, cyan and green, all on a black background, and a speaker should keep the same colour once assigned. By convention the narrator is yellow. Dashes are described as a legacy technique to be used with colour only when unavoidable, and non-human speakers get labels such as ROBOT: Hello, sir.

Ways to identify a speaker and what Sume supports, read 2026-10-03
MethodSourceOn Sume
Speaker colour (white, yellow, cyan, green)BBC 8.3, preferredOne text colour per caption job via design.colors; not per cue
Dash at the start of a new speaker's lineBBC 7.3, legacyPut it in the cue text
Label such as ROBOT:BBC 7.8Put it in the cue text
Sound and laughter cuesWCAG 1.2.2Put [laughs] or similar in the cue text
Automatic speaker labelsNot required by eitherNot exposed; STT 1.0 returns no speaker labels

Why can't Sume colour each speaker?

The caption job takes a design.colors block with base, active, stroke, accent, accent_deep and card, and those apply to the whole render (video captions). There is no per-cue colour field in the documented request, and design is unsupported on the punch and tiktok-green styles. The speech-to-text endpoint also has no public speaker labels, because diarization settings are fixed on the server; see the diarization comparison.

So the honest options are text-based. If you need true per-speaker colours, deliver a sidecar track instead of burning. The W3C WebVTT specification's sample interview marks each speaker with a voice span such as <v Roger Bingham> at the start of each cue (read 2026-10-03), which names the speaker in the file and lets a player decide how to show it. The sidecar post shows building one from a Sume transcript.

How do I mark speakers in burned-in cues?

Supply cues with text, start and end in seconds. Cues skip speech-to-text, so you provide the speaker turns, for example from your own notes or an editor's transcript. The script below adds a dash when the speaker changes and keeps a sound cue in brackets. It follows the BBC's dash style, which the BBC calls legacy; with a single colour on screen it is the clearest option left.

import json

turns = [  # (speaker, text, start, end) from your own diarized notes
    ("host", "Did you see the launch?", 0.0, 1.6),
    ("guest", "I did. It went smoothly.", 1.6, 3.4),
    ("guest", "Mostly.", 3.4, 4.0),
    ("host", "[laughs] Mostly?", 4.0, 5.2),
]

cues, prev = [], None
for who, text, start, end in turns:
    # BBC style: a dash starts each new speaker's line
    prefix = "- " if who != prev and prev is not None else ""
    cues.append({"text": prefix + text, "start": start, "end": end})
    prev = who

body = {
    "video_url": "https://media.sume.com/artifacts/example/interview.mp4",
    "style": "slam",
    "cues": cues,
    "design": {"colors": {"base": "#FFFFFF"}},
}
print(json.dumps(body, indent=2))

When is a dash enough, and when is it not?

A dash works for short two-person exchanges, where each new line is clearly a reply. It gets ambiguous in a long conversation with three or more voices, because a dash says only that the speaker changed, not who it is. There, prefix the first line of each turn with a name or label, as the BBC does for non-human speakers, and keep the label short so the card still fits the safe width.

A narrator or voice-over is the common case where colour is what viewers expect. Since a Sume job has one colour, one practical split is two renders of the same footage: dialogue clips in the default colour, and narration-only clips in a yellow set through design.colors.base. That keeps the BBC's narrator convention without per-cue colour, at the cost of one job per colour.

What should I check afterwards?

Run one clip and read it on a phone before you batch the rest.

  • Each dash starts a new speaker, and a continuing speaker's lines do not carry one.
  • Sound cues use brackets and describe the sound, not just the word laughter.
  • The text colour has enough contrast with the footage; the BBC puts its text on a black background.
  • Names or labels are consistent from the first card to the last.
  • The speaker turns come from a person or a diarizing tool, since Sume's STT will not tell you who spoke.

What does it cost?

A caption job is $0.20 per job for videos up to 60 seconds under the current fixed estimate, and cues avoid a second transcription; confirm live pricing in GET /v1/catalog. For sound cues in more depth, see the sound-cue checklist.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume