What ends a Sume STT sentence segment: . ! ? and the Japanese marks
Sume ends a segment on . ! ? 。 ! ? … plus optional trailing quotes or closing parentheses; the corner bracket 」 and a fullwidth ) are not on that list.

Sume's sentence segmentation ends a segment on a word token that finishes with one of . ! ? 。 ! ? …, optionally followed by closing quote marks or a closing parenthesis. A token ending in 。」 or ) does not match, so a Japanese line closed with a corner bracket stays open until the next qualifying token. This is the rule in the worker source, not a documented guarantee, so test it on your own output.
The rule, token by token
When you send segmentation: {"mode": "sentence"} with an STT request, the worker tests each timed word against one regular expression. The terminal characters are the ASCII ., !, ?, the ideographic 。, !, ? and the ellipsis …. After the terminal it allows one or more of ", ', ”, ’, ) or ]. Anything else after the terminal, such as 」 or the fullwidth ), stops the match.
The test is per token, not per line, so it depends on how the provider tokenizes the speech. A token that is only 。 still ends a segment; a token like ok.」 does not.
| Token | Ends a segment? | Why |
|---|---|---|
| done. | Yes | Terminal period |
| done?! | Yes | Last character is a terminal |
| "Yes." | Yes | Closing quote allowed after the period |
| (ok.) | Yes | Closing parenthesis allowed |
| 終わり。 | Yes | Ideographic full stop |
| 終わり。」 | No | Corner bracket is not in the closer set |
| ok.) | No | Fullwidth parenthesis is not in the closer set |
| ok, | No | Comma is not a terminal |
Reproduce it before you rely on it
The following check uses the same character classes. Python's \Z stands in for the end-of-string anchor.
import re
END = re.compile(r"[.!?\u3002\uff01\uff1f\u2026](?:[\"'\u201d\u2019)\]]+)?\Z")
for tok in ["done.", "done?!", "\u7d42\u308f\u308a\u3002",
"\u7d42\u308f\u308a\u3002\u300d", "ok.\uff09", "ok,"]:
print(repr(tok), bool(END.search(tok)))
What to do about it
If your content is Japanese with bracketed dialogue, expect some segments to run long. Merge or split on your side after you read segments[], or send the caption job cues that you built yourself. The video captions docs describe cues and segments input.
Limits: when no terminal punctuation appears at all, the whole text is one segment, and an unpunctuated tail is split on silences of 0.5 seconds or more. Punctuation is added by the provider, so which tokens carry it is not under your control.
A quick audit of your own results
Count how many of your segments end without a terminal, and look at the longest ones. A few unusually long segments in Japanese or Chinese text are the sign that a closing mark is blocking the match.
Also look at the shortest ones. A one-word segment is usually a token like an abbreviation or a lone mark that closed a sentence too early.
- Sort segments by
duration_secondsand read the top and bottom five. - Check where quoted speech starts and ends.
- Compare with the boundaries you would draw by hand for one minute of audio.
Sources
Related posts
More in Developers
- Which API key scope does each Sume webhook endpoint need?
The signing secret needs account:read, rotate and test deliveries need account:write, redeliver needs jobs:write or formats:write. Map each call to a key.
- Which Sume API requests count against the read rate limit?
Every GET and HEAD is a read, and so are POST /v1/generation/admission-preview and the MCP endpoint. Reads have their own per-key bucket, 40x the write one.
- Which Sume API routes are missing from the public OpenAPI document?
Asset upload routes, admission-preview, POST /v1/avatars, POST /v1/avatar-videos and the generic model runs path are implemented but hidden from OpenAPI.
- Which Sume API routes work without an API key?
Only health, catalog, openapi.json and the three bgm routes work without a key. Every other /v1 route needs a Bearer or x-api-key credential, never both.
Written by Sume