Sume STT silence split: it only runs after the last sentence end

Sume STT cuts at terminal punctuation first; the 0.5 s silence rule only splits the unpunctuated run after the last terminal, not earlier pauses.

4 min readSume
All posts

In Sume STT sentence mode, a pause of 0.5 seconds or more splits a segment only inside the unpunctuated run that is left after the last sentence-final mark. A long pause in the middle of a transcript that has punctuation later on does not create a boundary by itself, because those words are closed by the next terminal mark first. This is how the worker code is written, not a documented promise, so check it on your own clips.

How the grouping works

The worker walks the timed words in order and adds each to the current group. When a word ends with a terminal mark (period, exclamation, question mark, ideographic full stop or ellipsis, with optional closing quotes), the group is closed. Anything left when the words run out is the tail, and only the tail is passed to the silence splitter, with a default threshold of 0.5 seconds.

So there are two cases. A transcript with no terminal marks at all is all tail, and every 0.5 second pause splits it. A transcript with marks throughout has a short tail or none, and pauses inside sentences are ignored.

Where the silence rule applies (Sume worker source, repo read 2026-10-05)
TranscriptSegments come fromSilence rule used?
Fully punctuatedTerminal marksOnly on any short tail
No punctuation anywhereSilence gaps of 0.5 s or moreYes, everywhere
Punctuated start, unpunctuated endMarks, then silence on the endYes, on the end only
Punctuation only at the very endOne group up to the markNo

A small simulation

This replicates the grouping on a list of (word, start, end) tuples, so you can try your own transcripts without a job.

import re

END = re.compile(r"[.!?\u3002\uff01\uff1f\u2026][\"'\u201d\u2019)\]]*\Z")

def group(words, gap=0.5):
    out, cur = [], []
    for w in words:
        cur.append(w)
        if END.search(w[0]):
            out.append(cur); cur = []
    tail, last = [], None
    for w in cur:
        if last and w[1] - last[2] >= gap:
            out.append(tail); tail = []
        tail.append(w); last = w
    if tail:
        out.append(tail)
    return [" ".join(x[0] for x in g) for g in out]

print(group([("a", 0, .2), ("b", 1.5, 1.7), ("c.", 1.8, 2.0), ("d", 3, 3.2), ("e", 4, 4.2)]))

What to do with it

If your audio is a monologue where the provider punctuates inconsistently, do not expect the 0.5 second rule to rescue the early part. Either segment the words yourself from the words[] gaps, which are always returned, or send the segment list through a second pass that splits long ones at the largest internal gap.

If you only need pause positions, the word timings are enough: the gap between one word's end and the next word's start is the pause, and it needs no segmentation request at all.

Limits and how to verify

Whether a given transcript has punctuation in the middle depends on the provider, so the same audio can behave differently from one clip to the next. Run a sample with segmentation on, print the segment count, and compare it with the number of terminal marks in the returned text.

If the counts match and you wanted more cuts, the silence rule is not reaching those pauses, and a client-side split on the word gaps is the fix. The boundary_lead_ms setting, which defaults to 70 and ranges from 0 to 500, only moves where a boundary sits, not how many there are.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume