Korean TTS segment text has no spaces, unless a digit is in it
Sume's TTS segment text joins tokens with spaces only if one has a Latin letter or digit; else with nothing. Use segments for timing, your script for text.

In Sume TTS sentence segments, each segment's text is built from word tokens with a rule: if any token in the sentence contains a Latin letter or a digit, the tokens are joined with single spaces; if none does, they are joined with no separator at all. A Japanese line therefore reads naturally, but a Korean sentence made only of Hangul tokens can come back without its word spaces, while the same sentence with a number in it gets spaces. The result depends on what tokens the provider returns, so check yours.
The rule in the worker
The function is sentenceTextFromWords in the worker's TTS segmentation code. It tests the tokens against [A-Za-z0-9]. If any matches, it joins with a space; if none matches, it joins with an empty string. The reasoning is sound for Japanese and Chinese, which do not use spaces. It is less obviously right for Korean, which does.
The STT side is different. Its segment text always joins tokens with a single space, because STT tokens come from the provider's own transcript.
| Segment tokens | Joined with | Example output |
|---|---|---|
| Latin words | Space | Hello world |
| Hangul only | Nothing | Tokens run together |
| Hangul plus a digit | Space | Tokens spaced, digit included |
| Kana and kanji only | Nothing | Natural for Japanese |
| STT segments, any script | Space | Always spaced |
Reproduce the rule
The check below applies the same logic to a list of tokens, so you can see which of your sentences will be joined without spaces.
import re
def join_like_sume(tokens):
spaced = any(re.search(r"[A-Za-z0-9]", t) for t in tokens)
return (" " if spaced else "").join(tokens)
print(join_like_sume(["Hello", "world."]))
print(join_like_sume(["\uc548\ub155", "\ud558\uc138\uc694."]))
print(join_like_sume(["\uc548\ub155", "2026."]))
Limits
This is read from the worker code, not from a documented guarantee, and the provider may return tokens differently from what the rule assumes: a Korean token that already carries a trailing space, or a single token for a whole phrase, would not show the effect. The only way to know is to request timestamps.words and segmentation on a sample and read the text.
Segments also carry gapless time ranges, and the text join does not change them. Only the string is affected, which matters for subtitles, search and any comparison against your script. Use a wav or raw container if you also want per-segment audio slices; mp3 returns timings only.
What to do about it
Do not treat segment text as the canonical script. Keep your own source lines, and use the segment only for its start and end. If you generate captions from the segments, match each one to the line you sent by index, so the wording on screen is the wording you approved.
For Korean, the sibling rule matters too: the API requires at least one Hangul syllable when language is ko, and a Latin caption style rejects Hangul wording. See the stored posts on those two errors for the fixes.
Sources
Related posts
More in Developers
- TTS sentence slices: mp3 gives timings only, wav gives audio_urls
Sume TTS segmentation returns sentence timings for any container, but slice audio_urls only with wav or raw. Request shape, the 70 ms rule and when to pick wav.
- Pipes and @{} markers in a Sume TTS transcript: stripped, never spoken
Sume strips || cue breaks, @{...} markers and the display side of <display|spoken> before the voice reads; an empty result returns 400 transcript_no_speech.
- TTS with only an API key: list avatars and pick one with voice ready
You do not need a voice id for Sume TTS. List your avatars, pick one whose voice status is ready, and send its handle as avatar_handle. Python example.
- TTS word timings to karaoke captions: map words[] to text
Sume TTS returns words[] with start and end; captions take words as text, start, end. A short Python map skips speech-to-text and burns your exact script.
Written by Sume