Where to break subtitle lines: BBC rules and Sume caption cues

The BBC says one sentence per subtitle and no article-noun splits. Sume sentence segments cover the first rule; cues with a line break cover the rest.

5 min readSume
All posts

The BBC's subtitle guidelines say each subtitle should be a single complete sentence and that, when a line has to break, it should break at punctuation where possible and otherwise never split an article from its noun, a preposition from its phrase, a conjunction from its clause, a pronoun from its verb, or the parts of a complex verb (read 2026-10-03). Sume's sentence segmentation gives you the first half of that; the second half you can add yourself by authoring caption cues with a line break in the text.

Below: what the BBC rules say, what Sume does and does not do about them, and a short script that applies the break rule before burning.

What are the BBC's rules for where a subtitle breaks?

Section 3.2 recommends one complete sentence per subtitle, with exceptions depending on speech speed. Section 3.4 says to break subtitles and lines at logical points, ideally at a full stop, comma or dash, and lists the pairs not to split. Section 4.1 recommends 160 to 180 words per minute, which works out to a minimum on-screen time of around 0.3 seconds per word, for example 1.2 seconds for a four-word subtitle (read 2026-10-03). The BBC calls timing an editorial decision, so the 0.3 seconds is a target, not a hard limit.

BBC line-break and timing rules against what Sume does, read 2026-10-03
BBC ruleSectionWhat Sume does
One sentence per subtitle3.2Transcript segmentation mode sentence groups words on terminal punctuation
Break at punctuation3.4Sentence segments end at punctuation; inside a caption phrase, breaks follow word count, characters and pauses
Do not split article and noun, preposition and phrase, and so on3.4No grammar-aware break; supply your own line break through cues
About 0.3 s per word on screen4.1Follows the speaker's timing; check dense sentences yourself

What does Sume's sentence segmentation actually guarantee?

Ask for segmentation.mode: "sentence" on STT 1.0 or on a transcribed video inspect and you get gapless segments[]. Words are grouped on terminal punctuation, and an unpunctuated run is split on silence; the inspect route also accepts silence_split_seconds from 0.2 to 3 (video inspect). That satisfies one-sentence-per-subtitle in the usual case. Very long unpunctuated speech gets split by silence rather than by grammar, so a long ramble may still break mid-clause.

For the burned-in caption job, the phrase controls are phrasing.max_words (1 to 12), max_chars (4 to 60) and pause_seconds (0.05 to 3). These count words, characters and silence. They do not know that a pronoun belongs with its verb. For a deeper look at the character limits, see the Netflix 42-character post.

How do I apply the break rule before burning?

The caption endpoint accepts authored cues with text, start and end in seconds; the text may include a newline for a two-line card, and supplying cues skips speech-to-text (video captions). So take sentence segments, break each long one with a small rule, and send the result as cues. The script below breaks a sentence into two lines, preferring a comma, never ending a line on a word from a short avoid list, and balancing the line lengths. It is a heuristic for English, not a parser, so read the output.

Cues are phrase-level cards, so you give up the word-by-word timing that speech-aligned captions use. Pick this route when the break matters more than the animation.

import json

AVOID = {"a", "an", "the", "on", "in", "at", "of", "to", "for", "with", "about", "and",
         "but", "or", "he", "she", "it", "they", "we", "i", "you", "is", "are", "will", "have"}

def two_lines(text, max_chars=28):
    words = text.split()
    if len(text) <= max_chars:
        return text
    best = None
    for i in range(1, len(words)):
        last = words[i - 1]
        if last.lower() in AVOID:
            continue
        a, b = " ".join(words[:i]), " ".join(words[i:])
        score = abs(len(a) - len(b)) - (8 if last.endswith((",", ";", ":")) else 0)
        if best is None or score < best[0]:
            best = (score, a + "\n" + b)
    return best[1] if best else text

segments = [  # result["segments"] from a Sume STT job
    {"start": 0.0, "end": 3.4, "text": "If you only remember one thing, remember the first frame."},
    {"start": 3.4, "end": 6.0, "text": "Everything else is a bonus."},
]
cues = [{"text": two_lines(s["text"]), "start": s["start"], "end": s["end"]} for s in segments]
print(json.dumps({"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
                  "style": "slam", "cues": cues}, indent=2))

When is this worth the extra step?

Run it when readability is the product, and skip it when speed is. Whichever route you take, look at one rendered frame before batching: a heuristic that reads well on one sentence can leave an orphan word on the next, and burned-in text cannot be corrected by the viewer or by a platform player after upload.

  • Long, comma-heavy sentences on a vertical clip, where a bad break lands one word on the second line.
  • Educational or documentary clips where line breaks aid comprehension.
  • Corrected scripts: if you already have clean copy with its timings, cues skip the transcription charge for the caption job.
  • Not worth it for fast, short social captions where one to five words show at a time; there max_words does the job.
  • Not a substitute for proofreading: the segments are speech-to-text output.

What does it cost?

A caption job is $0.20 per job for videos up to 60 seconds under the current fixed estimate, whichever way it gets its text. The transcript that feeds the segments is billed at $0.01 per audio minute when you use STT or a transcribing inspect. Confirm live prices in GET /v1/catalog. If the dense sentences are the real problem, the reading-speed check flags which ones to break or shorten first.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume