Check caption reading speed from STT segments and flag dense sentences

Request sentence segments from Sume STT, compute characters per second for each, and flag the ones too dense to read on screen. Python and caption fixes.

4 min readSume
All posts

Submit the audio to Sume STT with segmentation.mode set to sentence, then divide each segment's character count by its duration. Sentences above your limit are the ones viewers will struggle to read as captions, so shorten them or loosen the caption phrasing before you burn anything in.

The limit is yours. The sketch uses 17 characters per second as an example house rule, not a platform requirement.

What do segments give me?

With segmentation.mode sentence, STT also returns segments[], each with index, text, start, end and duration_seconds. They are gapless (end equals the next start) and cover the audio you submitted. The docs say no sliced files are produced, only time ranges. boundary_lead_ms (0 to 500, default 70) controls how much of the pause goes to the previous sentence.

Segment fields used for the check (Sume OpenAPI, read 2026-10-02).
FieldUse
segments[].textCharacter count
segments[].startStart in seconds
segments[].endEnd in seconds
segments[].duration_secondsend - start

What does the check look like?

The sketch below flags the long sentence in a two-segment sample. It runs as written; feed it the real segments.

segments = [
    {"text": "Meet the bottle that keeps drinks cold.", "start": 0.0, "end": 2.6},
    {"text": "Twenty four hours of ice, no sweat on the outside, fits every cup holder.", "start": 2.6, "end": 5.0},
]
LIMIT = 17  # characters per second; your house rule, not a platform rule
for s in segments:
    cps = len(s["text"]) / (s["end"] - s["start"])
    flag = "DENSE" if cps > LIMIT else "ok"
    print(f"{cps:5.1f} cps  {flag}  {s['text'][:40]}")

How do I fix a dense sentence?

Standalone video captions already break speech into short phrases, and the design.phrasing group exposes max_words, max_chars and pause_seconds for styles that support design overrides (the docs say punch and tiktok-green do not). Lowering max_chars shows fewer characters at once. If that is not enough, pass your own cues with text, start and end, which burns authored copy without speech-to-text.

The cleanest fix is upstream: shorten the script so the narrator says fewer words in that sentence.

What are the limits of this check?

A sentence is a coarse unit. Captions show phrases, so a dense sentence may still read fine if it is broken into short cues, and a short sentence may be spoken fast. Characters per second also means different things in different scripts, so set a separate limit for Korean or Japanese rather than reusing an English number. Treat flagged rows as places to look at the rendered captions, not as errors.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume