Sume TTS without a language field: English default, Hangul fallback

What Sume TTS 1.0 does when you leave out language, why a Hangul-only script is the one fallback, and why you should set language yourself on every job.

4 min readSume
All posts

If you leave language out of a Sume TTS 1.0 request, the provider default is English. The one exception in the docs is a transcript written only in Hangul or kana, where Sume infers Korean or Japanese. Set language on every job anyway: it costs one field, and it turns a silent guess into a check.

What the field does

language is an optional code on POST /v1/tts-1.0/generate. When the chosen voice has a recorded language and your language differs, the API returns 409 tts_voice_language_mismatch instead of speaking with the wrong accent. Resend with confirm_language_mismatch: true if you mean it, for example a deliberate cross-language read.

How Sume TTS 1.0 picks a language - from the API reference (read 2026-10-07)
RequestResult
language set, matches voiceSpeaks in that language
language set, differs from voice409 tts_voice_language_mismatch unless confirm_language_mismatch is true
language omitted, mixed or Latin scriptEnglish at the provider
language omitted, Hangul-only or kana-only textKorean or Japanese is inferred

Where omission bites

MAI-Voice-2.1 and Gemini Flash TTS advertise dozens of languages, but that is a property of each vendor's model. On Sume, check the language of the chosen voice rather than assuming the model will adapt.

  • A Spanish script with no language field is read as English, with the accent that implies.
  • A Korean line with one English brand word is not Hangul-only, so the fallback does not apply.
  • A batch that mixes languages sends one wrong job and several right ones, which is hard to spot in review.

A request that sets it

Pass the code that matches both the script and the voice. If you do not know the voice's language, fetch the avatar first and read its voice record before you submit.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ko-greeting-001" \
  -d '{
    "transcript": "안녕하세요, 오늘의 소식을 전해 드립니다.",
    "avatar_handle": "anchor",
    "language": "ko"
  }'

What Sume does not do here

There is no per-sentence language switch in one job, and no automatic language detection beyond the narrow fallback above. For a bilingual script, split it into separate jobs with the right language on each, then join the audio on the timeline.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume