Does Dr. split a Sume STT segment? Abbreviations and sentence mode

Segmentation tests each token for a final period, so a token like Dr. or U.S. can end a segment early. Merge on your side with a small abbreviation list.

4 min readSume
All posts

It can. Sume's sentence segmentation tests every timed word for a sentence-final mark, and a period at the end of a token counts, so Dr. or U.S. followed by a name can close a segment in the middle of a sentence. This is an inference from the segmentation code rather than a documented behavior, and whether it occurs depends on whether the provider emits the period inside the token. The fix is a small merge step on your side.

Why this happens

The worker walks the words[] list and ends a segment whenever the current word matches the terminal pattern: a period, exclamation or question mark, an ideographic full stop or an ellipsis, optionally followed by a closing quote or parenthesis. It has no abbreviation list. So the token Dr. matches exactly as done. does.

The result is a segment that ends in Dr. and the next that starts with the surname. Timing stays gapless, because the boundary is min(lastWordEnd + lead, nextSentenceStart), so only the text split is wrong, not the audio coverage.

Likely false boundaries to check in your own results (rule from the Sume worker source, repo read 2026-10-05; examples are illustrative)
Token you may seeLooks like a stopReal sentence end
Dr.YesNo
Mr.YesNo
U.S.YesSometimes
etc.YesSometimes
No.YesSometimes

A merge step

Run this over the segments[] array: when a segment's text ends in a known abbreviation, join it with the next one and keep the first start and the last end. It works on plain dictionaries, so you can test it without any API call.

ABBR = ("Dr.", "Mr.", "Mrs.", "Ms.", "St.", "vs.")

def merge(segs):
    out = []
    carry = None
    for s in segs:
        if carry:
            s = {"text": carry["text"] + " " + s["text"],
                 "start": carry["start"], "end": s["end"]}
            carry = None
        if s["text"].endswith(ABBR):
            carry = s
        else:
            out.append(s)
    if carry:
        out.append(carry)
    return out

segs = [{"text": "Ask Dr.", "start": 0, "end": 1.0},
        {"text": "Lee today.", "start": 1.0, "end": 2.1}]
print(merge(segs))

Limits

Only some provider outputs will show this, and a list of abbreviations is never complete, so spot-check a few of your own transcripts before you add a merge. If captions are the goal, you can also skip segmentation: send the words to the captions endpoint as words or build cues yourself, since those inputs take whatever phrase boundaries you give them.

Do not merge across long silences. A segment that ends in St. followed by a two second gap is probably a real boundary, so check the gap between the two ends before you join.

How to tell whether you are affected

Do not assume the problem exists; check for it. Pull the segments[] for a few transcripts and look for text that ends in a capitalised abbreviation, then check what the next segment starts with.

If the next segment starts with a capitalised name or a lowercase word, you have found a false boundary. If you find none in a sample of ten files, leave segmentation alone and skip the merge step.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume