Where to break subtitle lines: BBC rules and Sume caption cues
The BBC says one sentence per subtitle and no article-noun splits. Sume sentence segments cover the first rule; cues with a line break cover the rest.

The BBC's subtitle guidelines say each subtitle should be a single complete sentence and that, when a line has to break, it should break at punctuation where possible and otherwise never split an article from its noun, a preposition from its phrase, a conjunction from its clause, a pronoun from its verb, or the parts of a complex verb (read 2026-10-03). Sume's sentence segmentation gives you the first half of that; the second half you can add yourself by authoring caption cues with a line break in the text.
Below: what the BBC rules say, what Sume does and does not do about them, and a short script that applies the break rule before burning.
What are the BBC's rules for where a subtitle breaks?
Section 3.2 recommends one complete sentence per subtitle, with exceptions depending on speech speed. Section 3.4 says to break subtitles and lines at logical points, ideally at a full stop, comma or dash, and lists the pairs not to split. Section 4.1 recommends 160 to 180 words per minute, which works out to a minimum on-screen time of around 0.3 seconds per word, for example 1.2 seconds for a four-word subtitle (read 2026-10-03). The BBC calls timing an editorial decision, so the 0.3 seconds is a target, not a hard limit.
| BBC rule | Section | What Sume does |
|---|---|---|
| One sentence per subtitle | 3.2 | Transcript segmentation mode sentence groups words on terminal punctuation |
| Break at punctuation | 3.4 | Sentence segments end at punctuation; inside a caption phrase, breaks follow word count, characters and pauses |
| Do not split article and noun, preposition and phrase, and so on | 3.4 | No grammar-aware break; supply your own line break through cues |
| About 0.3 s per word on screen | 4.1 | Follows the speaker's timing; check dense sentences yourself |
What does Sume's sentence segmentation actually guarantee?
Ask for segmentation.mode: "sentence" on STT 1.0 or on a transcribed video inspect and you get gapless segments[]. Words are grouped on terminal punctuation, and an unpunctuated run is split on silence; the inspect route also accepts silence_split_seconds from 0.2 to 3 (video inspect). That satisfies one-sentence-per-subtitle in the usual case. Very long unpunctuated speech gets split by silence rather than by grammar, so a long ramble may still break mid-clause.
For the burned-in caption job, the phrase controls are phrasing.max_words (1 to 12), max_chars (4 to 60) and pause_seconds (0.05 to 3). These count words, characters and silence. They do not know that a pronoun belongs with its verb. For a deeper look at the character limits, see the Netflix 42-character post.
How do I apply the break rule before burning?
The caption endpoint accepts authored cues with text, start and end in seconds; the text may include a newline for a two-line card, and supplying cues skips speech-to-text (video captions). So take sentence segments, break each long one with a small rule, and send the result as cues. The script below breaks a sentence into two lines, preferring a comma, never ending a line on a word from a short avoid list, and balancing the line lengths. It is a heuristic for English, not a parser, so read the output.
Cues are phrase-level cards, so you give up the word-by-word timing that speech-aligned captions use. Pick this route when the break matters more than the animation.
import json
AVOID = {"a", "an", "the", "on", "in", "at", "of", "to", "for", "with", "about", "and",
"but", "or", "he", "she", "it", "they", "we", "i", "you", "is", "are", "will", "have"}
def two_lines(text, max_chars=28):
words = text.split()
if len(text) <= max_chars:
return text
best = None
for i in range(1, len(words)):
last = words[i - 1]
if last.lower() in AVOID:
continue
a, b = " ".join(words[:i]), " ".join(words[i:])
score = abs(len(a) - len(b)) - (8 if last.endswith((",", ";", ":")) else 0)
if best is None or score < best[0]:
best = (score, a + "\n" + b)
return best[1] if best else text
segments = [ # result["segments"] from a Sume STT job
{"start": 0.0, "end": 3.4, "text": "If you only remember one thing, remember the first frame."},
{"start": 3.4, "end": 6.0, "text": "Everything else is a bonus."},
]
cues = [{"text": two_lines(s["text"]), "start": s["start"], "end": s["end"]} for s in segments]
print(json.dumps({"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"style": "slam", "cues": cues}, indent=2))
When is this worth the extra step?
Run it when readability is the product, and skip it when speed is. Whichever route you take, look at one rendered frame before batching: a heuristic that reads well on one sentence can leave an orphan word on the next, and burned-in text cannot be corrected by the viewer or by a platform player after upload.
- Long, comma-heavy sentences on a vertical clip, where a bad break lands one word on the second line.
- Educational or documentary clips where line breaks aid comprehension.
- Corrected scripts: if you already have clean copy with its timings, cues skip the transcription charge for the caption job.
- Not worth it for fast, short social captions where one to five words show at a time; there
max_wordsdoes the job. - Not a substitute for proofreading: the segments are speech-to-text output.
What does it cost?
A caption job is $0.20 per job for videos up to 60 seconds under the current fixed estimate, whichever way it gets its text. The transcript that feeds the segments is billed at $0.01 per audio minute when you use STT or a transcribing inspect. Confirm live prices in GET /v1/catalog. If the dense sentences are the real problem, the reading-speed check flags which ones to break or shorten first.
Sources
Related posts
More in Media tools
- Who is speaking in subtitles: BBC colours, dashes, and Sume cues
WCAG 1.2.2 wants speaker identification in captions. The BBC prefers colour, then dashes or labels. What Sume's burned-in captions can do and a dash script.
- YouTube Auto-sync captions vs Sume script_text: track or burned in
YouTube Auto-sync times your transcript into a caption track. Sume script_text aligns your wording into burned-in captions. What each needs and which to use.
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
Written by Sume