Korean voiceover: 1,200 Hangul characters cost 5.7 cents, so use NFC
Sume TTS bills $0.0475 per 1,000 transcript characters, so 1,200 Hangul characters cost 5.7 cents. Decomposed text can count two or three times more.

Sume TTS is priced at $0.0475 per 1,000 transcript characters, with spaces and punctuation counted, so a Korean script of 1,200 characters costs 1.2 x $0.0475 = $0.057, or 5.7 cents. The catch is that the same-looking Korean text can contain a very different number of characters. Hangul can be stored as one precomposed syllable or as two or three separate jamo, and text copied from some sources arrives in the decomposed form.
Per-character pricing and the 20,000-character request limit are from Sume's public catalog, read on 2026-10-09. How Sume counts decomposed text is not documented, so normalize before you submit.
Why the count can differ
Unicode has a composed form (NFC) in which one syllable is a single code point, and a decomposed form (NFD) in which it is split into its parts. A syllable with a final consonant is three jamo in NFD. If a counter sees code points, an NFD script of 1,200 visible syllables could count up to 3,600 characters, and the price would triple.
| Form | Code points per syllable | Characters for 1,200 syllables | Price |
|---|---|---|---|
| NFC (composed) | 1 | 1,200 | $0.057 |
| NFD, syllables with no final consonant | 2 | 2,400 | $0.114 |
| NFD, syllables with a final consonant | 3 | 3,600 | $0.171 |
A pre-flight check
Normalize the text to NFC in your code before you count it or submit it, and count after that step. In Python, unicodedata.normalize with the form NFC does it. Then compare len of the text with your budget, and keep it under 20,000.
Also set the language. Sume's guidance for Korean is to send a Hangul transcript with language set to ko; a Korean transcript without a language is mispronounced, and a voice-language mismatch returns a warning that you must confirm before the job is accepted. Never translate a Korean request into English text just to save characters.
import unicodedata
def tts_chars(text):
clean = unicodedata.normalize("NFC", text)
return clean, len(clean)
nfd = unicodedata.normalize("NFD", "안녕하세요 여러분")
clean, n = tts_chars(nfd)
print(len(nfd), n, round(n / 1000 * 0.0475, 5))Budget examples
A 600-character caption script is $0.0285. A 20,000-character script, the largest single request, is $0.95. These numbers assume NFC input and the catalog rate; the TTS tool in Sume's hosted MCP server has a dry_run option that previews cost, so test a short script first and compare the preview with your own count.
- Normalize to NFC, then count.
- Send language ko with a Hangul transcript.
- Stay under 20,000 characters per request.
Where decomposed text comes from
Decomposed Hangul shows up when text is copied from certain file systems, some older documents, and some editors that store filenames or form input in the decomposed form. It looks identical on screen, so you cannot see the problem by eye. Only a character count or a normalization call reveals it.
The fix costs nothing. Add the normalization step to the same function that strips stray whitespace and checks the 20,000-character limit, and every script passes through it. Then the cost you estimate from the Hangul you see is the cost you pay.
Sources
Related posts
More in Developers
- Lint a video against TikTok, Reels, Feed and Shorts limits in Python
A 24-line Python check for duration, file size and average bitrate against the limits TikTok, Meta and YouTube publish, plus where Sume jobs fit before it.
- Lint a Video Router request offline: reference limits in Python
Check reference_image_urls (10), reference_video_urls (3), reference_audio_urls (1-5) and the video_url rule in Python before POST /v1/video-router/generate.
- List Sume's video endpoints from the OpenAPI JSON in Python
Pull the OpenAPI snapshot, print every /v1/videos and Video Router operation, and know which copy is the source of truth. A short Python script.
- List Sume video model limits with curl, jq and column in one table
curl GET /v1/videos/models into jq and column to see each model's duration range, resolutions and audio flag in one table. Documented limits for six models.
Written by Sume