Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.

Pocket TTS speaks English, French, German, Portuguese, Italian and Spanish according to Kyutai's model card and its 2026-05-04 blog post. The GitHub README adds Dutch as a seventh. I read all three on 2026-10-03, so the honest answer is six confirmed by two Kyutai pages, plus Dutch listed by one.
No Korean, Japanese, Chinese or Hindi appears on any of those pages. If you need those, a self-hosted Pocket TTS is not the route today, and Sume's TTS language field is the one to look at.
What exactly does each Kyutai page list?
The pages were written at different times, which probably explains the mismatch: the blog post that announced the multilingual release is dated 2026-05-04, and the README may have moved since.
| Language | README | Model card | Blog (2026-05-04) |
|---|---|---|---|
| English | Listed | Listed | Listed |
| French | Listed | Listed | Listed |
| German | Listed | Listed | Listed |
| Portuguese | Listed | Listed | Listed |
| Italian | Listed | Listed | Listed |
| Spanish | Listed | Listed | Listed |
| Dutch | Listed | Not listed | Not listed |
How does Sume TTS pick a language?
Sume's TTS request has a language field: the language the voice speaks the transcript in, as a BCP-47 or ISO-639 code such as ko, ja or en. The schema says to set it for every non-English transcript, because an omitted value defaults to English at the provider. As a fallback Sume infers Korean or Japanese from a transcript that is only Hangul or kana, but the field is the reliable way.
There is also a guard rail: if the voice you picked does not match the language, the API returns a mismatch warning, and you resend with confirm_language_mismatch: true only after a person has agreed. The request never translates a transcript into English for you.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "Votre commande est partie aujourd'"'"'hui.",
"avatar_id": "'"$SUME_AVATAR_ID"'",
"language": "fr",
"mode": "sync",
"wait_timeout_seconds": 30
}'
What should you check before you rely on a language?
A language list says the model can speak it, not that every voice does it well. Before you ship a campaign in a language, run the same two sentences through the voice you plan to use and listen.
- Generate a short sample in each target language with the real voice, not a demo voice.
- For Sume, set
languageon every non-English request and keep the transcript in that language; the field does not translate. - For Pocket TTS, test Dutch separately, since only the README lists it.
- Round-trip the output through speech-to-text and compare the words; Check a TTS take with an STT round trip shows the script.
How do you handle a mixed-language script?
A script that switches language mid-sentence is the case where language lists mislead most. A voice set to one language may read a foreign brand name in that language's sounds, which sometimes works and sometimes sounds wrong. The reliable approach is to split the script by language and synthesise each part with the right setting, then join the clips.
With Sume that is several TTS jobs, each with its own language, followed by one POST /v1/timeline-1.0/audio concat of up to 20 parts, which joins them in the sample domain with no re-synthesis and no silence at the seams. With a self-hosted Pocket TTS you would choose the language per call, using whichever of the six or seven languages the model supports for each part.
Plan the join in advance: keep every clip in the same channel layout (the concat refuses parts that differ), and render to wav if the joined file will feed a lip-sync or avatar step, because mp3 re-adds encoder padding at each edge. Join voiceover takes into one track has the full request. The README, the model card and the May blog post differ, and I could not tell from the text which is newest. Re-read the Kyutai blog before you plan around a list, because the project is moving; its most recent Pocket TTS post, dated 2026-09-28, was about the sampler, not languages.
Where does that leave a multilingual project?
For six European languages on your own CPU, Pocket TTS is a legitimate option. For a catalogue that includes Korean or Japanese, or when you want the audio file hosted and billed in one place, a Sume job with language set is the simpler path. Build an AI dubbing pipeline from STT, translate and TTS puts the language field to work end to end.
Sources
Related posts
More in Models
- Omni Flash vs H3 Max for reference-to-video: limits side by side
Gemini Omni Flash 1.1 takes 10 images and 3 short videos; MiniMax H3 Max adds audio refs, 12 files total. Limits, tags and a sample request on Sume.
- Shortest AI video clip via API: 2 seconds on Wan 3.0 only
Sume's video catalog by minimum duration: Wan 3.0 takes 2 seconds, Omni 3, Seedance and Kling 4, MiniMax H3 5. Plus how to trim below a floor.
- Silent AI video: generate_audio false or drop the audio after
Seedance, Kling and Wan take generate_audio false; Omni and MiniMax H3 always make sound. How to get a silent clip on Sume, and what it costs.
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
Written by Sume