Korean TTS segment text has no spaces, unless a digit is in it

Sume's TTS segment text joins tokens with spaces only if one has a Latin letter or digit; else with nothing. Use segments for timing, your script for text.

4 min readSume
All posts

In Sume TTS sentence segments, each segment's text is built from word tokens with a rule: if any token in the sentence contains a Latin letter or a digit, the tokens are joined with single spaces; if none does, they are joined with no separator at all. A Japanese line therefore reads naturally, but a Korean sentence made only of Hangul tokens can come back without its word spaces, while the same sentence with a number in it gets spaces. The result depends on what tokens the provider returns, so check yours.

The rule in the worker

The function is sentenceTextFromWords in the worker's TTS segmentation code. It tests the tokens against [A-Za-z0-9]. If any matches, it joins with a space; if none matches, it joins with an empty string. The reasoning is sound for Japanese and Chinese, which do not use spaces. It is less obviously right for Korean, which does.

The STT side is different. Its segment text always joins tokens with a single space, because STT tokens come from the provider's own transcript.

How segment text is joined (Sume worker source, repo read 2026-10-05)
Segment tokensJoined withExample output
Latin wordsSpaceHello world
Hangul onlyNothingTokens run together
Hangul plus a digitSpaceTokens spaced, digit included
Kana and kanji onlyNothingNatural for Japanese
STT segments, any scriptSpaceAlways spaced

Reproduce the rule

The check below applies the same logic to a list of tokens, so you can see which of your sentences will be joined without spaces.

import re

def join_like_sume(tokens):
    spaced = any(re.search(r"[A-Za-z0-9]", t) for t in tokens)
    return (" " if spaced else "").join(tokens)

print(join_like_sume(["Hello", "world."]))
print(join_like_sume(["\uc548\ub155", "\ud558\uc138\uc694."]))
print(join_like_sume(["\uc548\ub155", "2026."]))

Limits

This is read from the worker code, not from a documented guarantee, and the provider may return tokens differently from what the rule assumes: a Korean token that already carries a trailing space, or a single token for a whole phrase, would not show the effect. The only way to know is to request timestamps.words and segmentation on a sample and read the text.

Segments also carry gapless time ranges, and the text join does not change them. Only the string is affected, which matters for subtitles, search and any comparison against your script. Use a wav or raw container if you also want per-segment audio slices; mp3 returns timings only.

What to do about it

Do not treat segment text as the canonical script. Keep your own source lines, and use the segment only for its start and end. If you generate captions from the segments, match each one to the line you sent by index, so the wording on screen is the wording you approved.

For Korean, the sibling rule matters too: the API requires at least one Hangul syllable when language is ko, and a Latin caption style rejects Hangul wording. See the stored posts on those two errors for the fixes.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume