What ends a Sume STT sentence segment: . ! ? and the Japanese marks

Sume ends a segment on . ! ? 。 ! ? … plus optional trailing quotes or closing parentheses; the corner bracket 」 and a fullwidth ) are not on that list.

4 min readSume
All posts

Sume's sentence segmentation ends a segment on a word token that finishes with one of . ! ? 。 ! ? …, optionally followed by closing quote marks or a closing parenthesis. A token ending in 。」 or ) does not match, so a Japanese line closed with a corner bracket stays open until the next qualifying token. This is the rule in the worker source, not a documented guarantee, so test it on your own output.

The rule, token by token

When you send segmentation: {"mode": "sentence"} with an STT request, the worker tests each timed word against one regular expression. The terminal characters are the ASCII ., !, ?, the ideographic 。, !, ? and the ellipsis …. After the terminal it allows one or more of ", ', ”, ’, ) or ]. Anything else after the terminal, such as 」 or the fullwidth ), stops the match.

The test is per token, not per line, so it depends on how the provider tokenizes the speech. A token that is only 。 still ends a segment; a token like ok.」 does not.

Which tokens end a segment (regular expression in the Sume worker, checked in Python, repo read 2026-10-05)
TokenEnds a segment?Why
done.YesTerminal period
done?!YesLast character is a terminal
"Yes."YesClosing quote allowed after the period
(ok.)YesClosing parenthesis allowed
終わり。YesIdeographic full stop
終わり。」NoCorner bracket is not in the closer set
ok.)NoFullwidth parenthesis is not in the closer set
ok,NoComma is not a terminal

Reproduce it before you rely on it

The following check uses the same character classes. Python's \Z stands in for the end-of-string anchor.

import re

END = re.compile(r"[.!?\u3002\uff01\uff1f\u2026](?:[\"'\u201d\u2019)\]]+)?\Z")

for tok in ["done.", "done?!", "\u7d42\u308f\u308a\u3002",
            "\u7d42\u308f\u308a\u3002\u300d", "ok.\uff09", "ok,"]:
    print(repr(tok), bool(END.search(tok)))

What to do about it

If your content is Japanese with bracketed dialogue, expect some segments to run long. Merge or split on your side after you read segments[], or send the caption job cues that you built yourself. The video captions docs describe cues and segments input.

Limits: when no terminal punctuation appears at all, the whole text is one segment, and an unpunctuated tail is split on silences of 0.5 seconds or more. Punctuation is added by the provider, so which tokens carry it is not under your control.

A quick audit of your own results

Count how many of your segments end without a terminal, and look at the longest ones. A few unusually long segments in Japanese or Chinese text are the sign that a closing mark is blocking the match.

Also look at the shortest ones. A one-word segment is usually a token like an abbreviation or a lone mark that closed a sentence too early.

  • Sort segments by duration_seconds and read the top and bottom five.
  • Check where quoted speech starts and ends.
  • Compare with the boundaries you would draw by hand for one minute of audio.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume