Text to speech Korean API: 90+ languages claim vs what Sume lists
Eleven v4 says 90+ languages. Sume's TTS schema publishes no language count, only a language field and five Sonic ids. Request Korean or Japanese and test it.

For a text to speech Korean API on Sume, send language: "ko" with your Korean transcript to POST /v1/tts-1.0/generate, or pass the same field to the TTS Router with one of five Sonic ids. ElevenLabs says its new Eleven v4 covers over 90 languages; Sume publishes no such count, so the way to know a language works is to render a sentence and listen.
What does ElevenLabs claim about languages?
ElevenLabs' September 28, 2026 announcement says both Eleven v4 and Eleven v4 Turbo support over 90 languages. That is the vendor's claim for its own models. Sume's router does not list an Eleven id, so the number does not describe what you get through Sume.
What does Sume actually publish?
Two things describe language on Sume's side, and neither is a list.
| Field | What the schema says |
|---|---|
language | 2 to 16 characters, BCP-47 or ISO-639, for example ko, ja, en |
| Default | English at the provider when omitted; ko or ja is inferred from a Hangul- or kana-only transcript |
Router model | One of sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview |
| Language count | Not published |
How do I request Korean or Japanese?
Set language explicitly rather than relying on inference. Use a voice you own: an avatar_id or avatar_handle whose voice is ready, or a voice.id. The router takes the same body plus a required model.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "sonic-3.6",
"transcript": "안녕하세요. 새 대시보드를 소개합니다.",
"avatar_handle": "product_host",
"language": "ko"
}'Can I ask Sume for an Eleven voice or model id?
No. The router's model field is an enum of the five Sonic ids, and the catalog describes version 1 as Sonic only. A voice must be a TTS voice UUID or a Sume Voices library id starting with voi_; a voice name from another TTS ecosystem is rejected with 400 before a job is queued or credits are reserved. So Eleven-style inline tags and voice names have no place in a Sume request; write plain text and pick the delivery with the schema's generation_config.
If Korean is your only need, the smaller question is not the language count at all. It is whether the voice you picked reads your Korean naturally, which only a listening test answers.
How do I check a language before I commit?
Render one short sentence per target language, at the exact voice you plan to use, and listen. A voice with a recorded language different from the requested one can be flagged: the schema mentions a language-mismatch warning that you confirm with confirm_language_mismatch. Route requests are billed by character, so a short test costs little. For a longer walk through Korean, see Korean text to speech API.
Sources
Related posts
More in Models
- Eleven v4 audio tags [laughs]: how to use them, and Sume's controls
Eleven v4 reads inline tags such as [laughs] and [light rain] in the text. Sume's TTS schema has no tag syntax; it has speed, volume and emotion controls.
- FLUX.2 dev commercial use: what the model card says
FLUX.2 [dev] weights carry the FLUX non-commercial license, yet the card says outputs can be used commercially as that license describes. The exact wording.
- FLUX.2 flex text rendering: prompt tips and the Sume model id
Black Forest Labs positions FLUX.2 [flex] for text and fine detail. How its typography guidance reads, and how to call black-forest-labs/flux.2-flex on Sume.
- FLUX.2 hex color prompt: how to write it, with an API call
Black Forest Labs says FLUX.2 matches hex codes in prompts. The syntax it documents, the limits it admits, and a flux.2-pro call on Sume.
Written by Sume