TTS language omitted: Spanish text comes out with an English default
On Sume TTS, an omitted language defaults to English at the provider, with a ko or ja fallback only for Hangul or kana text. Set language for all others.

If you leave language off a Sume TTS request, the voice is told the text is English. The OpenAPI description for the field says to set it for every non-English transcript, because an omitted value defaults to English at the provider, and that Sume infers Korean or Japanese only as a fallback from a transcript made of Hangul or kana alone. A Spanish, French or German script gets no such rescue: send language explicitly.
This applies to both POST /v1/tts-1.0/generate and POST /v1/tts-router/generate; both schemas carry the same description. It was read from the OpenAPI file in the Sume docs on 2026-10-02.
What does the schema say about the field?
language takes 2 to 16 characters and is described as a BCP-47 or ISO-639 code such as ko, ja or en. The final sentence of the description is a rule for callers and agents alike: never translate a non-English request into English.
That sentence matters for dubbing, where text moves between languages. The TTS call speaks the text it receives; it does not translate it.
| Case | What the schema says |
|---|---|
| Field omitted, English text | English, the default |
| Field omitted, Hangul or kana only | Sume infers ko or ja as a fallback |
| Field omitted, Spanish or other Latin-script text | No inference stated; the default is English |
| Field set | The language the voice speaks the transcript in |
How should I set it in a dubbing pipeline?
Carry the target language code next to each translated line and pass it on every TTS request. If one batch holds several languages, send one request per language; the multilingual text to speech guide covers that shape.
For the source side, speech-to-text takes an optional language_code hint and auto-detects when you omit it. Caption jobs follow the same split: language is a speech-to-text hint there, and it does not choose the style or the font.
What should I check before a batch?
Run one short line in each target language and listen for English-sounding pronunciation, which is the sign the language field did not arrive. Keep the code in a single variable so the request body and your file naming agree. The Sume docs do not list which languages each voice supports, so a listening check is the only proof for a given pair.
Sources
Related posts
More in Developers
- TTS sentence slices have no audio_url: emit_audio needs wav or raw
With segmentation on, Sume TTS only returns a sample-exact audio_url per sentence for wav or raw output. For mp3 you get timings and a warning instead.
- TTS speed: slow, normal, fast or generation_config.speed 0.6 to 1.5?
Sume TTS marks the slow, normal and fast speed enum deprecated. Send generation_config.speed from 0.6 to 1.5, plus volume 0.5 to 2 and an emotion string.
- TTS call returned processing in sync mode: poll, do not resubmit
Sume TTS in sync mode waits at most 30 seconds. If the job is not terminal, poll status_url, and retry a submit only with the same Idempotency-Key.
- Test call audio for voice agents: Sume TTS at 8 kHz mu-law
Generate repeatable phone-quality test utterances for a voice agent with Sume TTS output_format: 8000 Hz, pcm_mulaw. Fields, limits and a runnable script.
Written by Sume