Sume STT silence split: it only runs after the last sentence end
Sume STT cuts at terminal punctuation first; the 0.5 s silence rule only splits the unpunctuated run after the last terminal, not earlier pauses.

In Sume STT sentence mode, a pause of 0.5 seconds or more splits a segment only inside the unpunctuated run that is left after the last sentence-final mark. A long pause in the middle of a transcript that has punctuation later on does not create a boundary by itself, because those words are closed by the next terminal mark first. This is how the worker code is written, not a documented promise, so check it on your own clips.
How the grouping works
The worker walks the timed words in order and adds each to the current group. When a word ends with a terminal mark (period, exclamation, question mark, ideographic full stop or ellipsis, with optional closing quotes), the group is closed. Anything left when the words run out is the tail, and only the tail is passed to the silence splitter, with a default threshold of 0.5 seconds.
So there are two cases. A transcript with no terminal marks at all is all tail, and every 0.5 second pause splits it. A transcript with marks throughout has a short tail or none, and pauses inside sentences are ignored.
| Transcript | Segments come from | Silence rule used? |
|---|---|---|
| Fully punctuated | Terminal marks | Only on any short tail |
| No punctuation anywhere | Silence gaps of 0.5 s or more | Yes, everywhere |
| Punctuated start, unpunctuated end | Marks, then silence on the end | Yes, on the end only |
| Punctuation only at the very end | One group up to the mark | No |
A small simulation
This replicates the grouping on a list of (word, start, end) tuples, so you can try your own transcripts without a job.
import re
END = re.compile(r"[.!?\u3002\uff01\uff1f\u2026][\"'\u201d\u2019)\]]*\Z")
def group(words, gap=0.5):
out, cur = [], []
for w in words:
cur.append(w)
if END.search(w[0]):
out.append(cur); cur = []
tail, last = [], None
for w in cur:
if last and w[1] - last[2] >= gap:
out.append(tail); tail = []
tail.append(w); last = w
if tail:
out.append(tail)
return [" ".join(x[0] for x in g) for g in out]
print(group([("a", 0, .2), ("b", 1.5, 1.7), ("c.", 1.8, 2.0), ("d", 3, 3.2), ("e", 4, 4.2)]))
What to do with it
If your audio is a monologue where the provider punctuates inconsistently, do not expect the 0.5 second rule to rescue the early part. Either segment the words yourself from the words[] gaps, which are always returned, or send the segment list through a second pass that splits long ones at the largest internal gap.
If you only need pause positions, the word timings are enough: the gap between one word's end and the next word's start is the pause, and it needs no segmentation request at all.
Limits and how to verify
Whether a given transcript has punctuation in the middle depends on the provider, so the same audio can behave differently from one clip to the next. Run a sample with segmentation on, print the segment count, and compare it with the number of terminal marks in the returned text.
If the counts match and you wanted more cuts, the silence rule is not reaching those pauses, and a client-side split on the word gaps is the fix. The boundary_lead_ms setting, which defaults to 70 and ranges from 0 to 500, only moves where a boundary sits, not how many there are.
Sources
Related posts
More in Developers
- Sume sync mode: set your client timeout above 30 s, then poll
Sume sync waits at most 30 seconds. If your HTTP client gives up at 30 s too, you lose the job id. Use a 45 s timeout, then poll and reuse the Idempotency-Key.
- Sume TTS 1.0 or TTS Router: which endpoint for a voiceover?
Both bill $0.0475 per 1,000 characters. TTS 1.0 always runs sonic-3.6; the router makes you name a Sonic id. Pick by whether you need to pin the engine.
- Sume TTS 400: send transcript or transcript_source, never both
A TTS request needs exactly one of transcript or transcript_source. Both, or neither, is an error. Live-commerce Formats need the source. A validator.
- Sume TTS has no SSML field: speed, emotion, pronunciation dictionary
Moving an SSML voice script to Sume TTS? The request takes plain transcript text, so use speed 0.6 to 1.5, volume, emotion and a pronunciation dictionary id.
Written by Sume