Eleven v4 claims 90+ languages: check Korean before you ship
A 90+ language claim is not a pronunciation guarantee. How to QA Korean voiceover on Eleven v4, and what Sume's voice-language guard and receipts prove.

Treat 90+ languages as a list of languages the model accepts, not as proof that Korean ad copy will sound right. ElevenLabs states the figure for Eleven v4 (model id eleven_v4), and its announcement also lists IPA phoneme support and a pronunciation dictionary. None of that replaces a listening pass by a native speaker on your own product names, numbers and prices.
The pages read give no per-language quality figure for Korean, so this post gives a QA routine rather than a verdict.
What the vendor pages state
| Item | Stated |
|---|---|
| Languages | 90+ |
| Model id | eleven_v4 |
| Pronunciation controls | Better IPA phoneme support; pronunciation dictionary |
| Per-generation limit | Up to 10,000 characters |
| Voice Design and cloning | Voice Design; Instant Voice Clones from 10 seconds of audio |
| Blind-test claim | About 75% listener preference (vendor claim) |
A five-step Korean QA routine
Run this on a fixed test script before any batch render. Keep the script and settings so the next model release can be compared on identical input.
- Build a 10-line script with the hard cases: brand names in Latin letters, prices and phone numbers, dates, a loanword, and a long sentence.
- Render it with each candidate voice and settings, and keep every output.
- Have a native speaker mark each line: correct, odd stress, or wrong reading. Count the lines, do not average a feeling.
- Fix reading problems with the pronunciation dictionary or phoneme input rather than by respelling the copy.
- Re-run the same script after any model or voice change.
What Sume checks, and what it does not
Sume does not ship Eleven v4; its TTS Router serves Cartesia Sonic models only. The language handling it does add sits before the audio is made. Each stored voice has a primary language, set when the voice is created. If the language on your request disagrees with it, the API returns HTTP 409 tts_voice_language_mismatch before any job is created or charge made. You confirm by retrying the same request with confirm_language_mismatch: true.
For script-bound jobs, the job also carries a receipt with the submitted text hash. The Sume TTS reference is explicit that this proves the submitted text, not the pronunciation, so listening QA stays a separate step on any engine.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ko-check-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Test line for review.",
"voice": { "id": "YOUR_VOICE_UUID" },
"language": "ko"
}'The request above omits confirm_language_mismatch on purpose. If the voice's primary language is not Korean, the API answers 409; add confirm_language_mismatch: true to a retry only after you have asked a person; the guard exists so a wrong voice does not reach a paid render silently. Drop the field when the voice language already matches.
Sources
Related posts
More in Use cases
- 10 seconds to clone a voice: ElevenLabs IVC vs cloning in Sume
ElevenLabs Instant Voice Clones start from 10 seconds of audio. Sume voice cloning is free of charge and sets a language per voice; here is how they differ.
- Email header image 600x200: render 1920x640 at exactly 3:1
A 600x200 email header is 3:1, the widest shape GPT Image 2.5 accepts. Render 1920x640, then export 1200x400 for retina and 600x200 for 1x with Pillow.
- Employee handbook page to a 20 s training clip: one policy, one scene
Turn one handbook page into a Wan 3.0 training clip on Sume: pick the policy, picture the behavior, type the rule as captions, and review it with HR before use.
- Employee onboarding videos: ten 30-second Wan 3.0 modules for $37.50
Ten 30 second onboarding clips on Sume's Wan 3.0 cost $37.50 at 720p, $18.75 at 480p or $75.00 at 1080p. Concurrency by plan and what to keep out of frames.
Written by Sume