Grok TTS language "auto" vs Sume's explicit language field
xAI TTS can auto-detect the language of your text. Sume TTS cannot: omit language and Spanish is read as English, except Korean and Japanese.

Does Sume detect the language of a TTS script like Grok's "auto"?
No. xAI lists "auto" among its language options, so the service works out the language of your text. Sume TTS 1.0 asks you to name it: the language field takes a BCP-47 or ISO-639 code such as ko, ja or en, and the contract says to set it for every non-English transcript.
If you leave it out, the OpenAPI description says the request defaults to English at the provider. Sume adds one fallback: it infers ko or ja from a transcript that is only Hangul or only kana.
What happens to a Spanish script with no language set?
It is spoken with English pronunciation rules. The words are the same, so the job completes and is billed, and you hear an accent mismatch instead of an error. That is the main risk of moving from an auto-detecting API to an explicit one.
Other languages are not in the fallback. Hindi, Arabic, French and every Latin-script language other than English need an explicit code.
| Case | xAI TTS | Sume TTS 1.0 |
|---|---|---|
| Auto-detect | Yes, "auto" option | No |
| Language omitted | Not described on the page read | English at the provider |
| Korean or Japanese only text | Auto handles it | Inferred as ko or ja |
| Spanish text, no code | Auto handles it | Read as English |
| Voice and language disagree | All voices speak every supported language | 409 or a confirmation step before any job |
How do you make your code set the language every time?
Make language a required argument in your own wrapper, take it from the content record, and send it on every request. A request whose voice has a different primary language returns HTTP 409 tts_voice_language_mismatch, or a non-error warning in the MCP tool that you confirm and retry with the same idempotency key. The check happens before any job is created or credit reserved.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: es-welcome-001" \
-d '{
"transcript": "Bienvenidos a nuestra tienda. Hoy tenemos ofertas.",
"avatar_handle": "narrator",
"language": "es"
}'How do you detect language yourself before sending?
If your pipeline receives mixed-language text, detect in your own code or in the system that wrote the copy. A translation step usually knows the target language already, so pass that value through rather than guessing from text. Mixed-language sentences are the hard case; xAI does not describe how "auto" resolves them in the page read, and Sume reads the whole request in the one language you give.
Split code-switched copy into one job per language run, set each language, and join the results with timeline audio. It costs a few more jobs and removes the guesswork.
What are the limits of this advice?
The xAI page lists 20 languages plus auto in the part we read. Sume does not publish a language list in the TTS contract, so test each language you plan to ship with a short sample before a batch.
The related posts cover Hinglish and Odia and Urdu codes, which are the cases people most often get wrong.
Sources
Related posts
More in Comparisons
- Grok TTS speech tags like laugh and whisper vs Sume's emotion field
xAI lets you write [laugh] and wrap text in whisper or singing tags. Sume TTS documents no tag grammar, only emotion, speed and volume in generation_config.
- H3 Max Recast vs Genjutsu: which person swap to call on Sume
Sume lists two person-swap video rows. Recast: 1-4 people in a 5-30 s clip at 768p or 1080p. Genjutsu: 1-8 images at 480p or 720p. How to choose.
- Hedra 402 INSUFFICIENT_BALANCE vs Sume insufficient_credits
Hedra returns 402 INSUFFICIENT_BALANCE until you add funds; Sume returns 402 insufficient_credits. How each wallet check behaves and how to preflight a job.
- Hedra Avatar needs start frame and audio; Sume takes a handle and scri
Hedra Avatar generates from a start frame plus an audio track, up to 10 minutes. Sume's talking-video takes an avatar handle and a script, up to 60 seconds.
Written by Sume