Text to speech Korean API: 90+ languages claim vs what Sume lists

Eleven v4 says 90+ languages. Sume's TTS schema publishes no language count, only a language field and five Sonic ids. Request Korean or Japanese and test it.

4 min readSume
All posts

For a text to speech Korean API on Sume, send language: "ko" with your Korean transcript to POST /v1/tts-1.0/generate, or pass the same field to the TTS Router with one of five Sonic ids. ElevenLabs says its new Eleven v4 covers over 90 languages; Sume publishes no such count, so the way to know a language works is to render a sentence and listen.

What does ElevenLabs claim about languages?

ElevenLabs' September 28, 2026 announcement says both Eleven v4 and Eleven v4 Turbo support over 90 languages. That is the vendor's claim for its own models. Sume's router does not list an Eleven id, so the number does not describe what you get through Sume.

What does Sume actually publish?

Two things describe language on Sume's side, and neither is a list.

Language-related fields in Sume's TTS schema, read 2026-09-29.
FieldWhat the schema says
language2 to 16 characters, BCP-47 or ISO-639, for example ko, ja, en
DefaultEnglish at the provider when omitted; ko or ja is inferred from a Hangul- or kana-only transcript
Router modelOne of sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview
Language countNot published

How do I request Korean or Japanese?

Set language explicitly rather than relying on inference. Use a voice you own: an avatar_id or avatar_handle whose voice is ready, or a voice.id. The router takes the same body plus a required model.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "안녕하세요. 새 대시보드를 소개합니다.",
    "avatar_handle": "product_host",
    "language": "ko"
  }'

Can I ask Sume for an Eleven voice or model id?

No. The router's model field is an enum of the five Sonic ids, and the catalog describes version 1 as Sonic only. A voice must be a TTS voice UUID or a Sume Voices library id starting with voi_; a voice name from another TTS ecosystem is rejected with 400 before a job is queued or credits are reserved. So Eleven-style inline tags and voice names have no place in a Sume request; write plain text and pick the delivery with the schema's generation_config.

If Korean is your only need, the smaller question is not the language count at all. It is whether the voice you picked reads your Korean naturally, which only a listening test answers.

How do I check a language before I commit?

Render one short sentence per target language, at the exact voice you plan to use, and listen. A voice with a recorded language different from the requested one can be flagged: the schema mentions a language-mismatch warning that you confirm with confirm_language_mismatch. Route requests are billed by character, so a short test costs little. For a longer walk through Korean, see Korean text to speech API.

Sources

Related posts

More in Models

All Models posts

Written by Sume