Pocket TTS languages: six or seven, and Sume's language field

Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.

5 min readSume
All posts

Pocket TTS speaks English, French, German, Portuguese, Italian and Spanish according to Kyutai's model card and its 2026-05-04 blog post. The GitHub README adds Dutch as a seventh. I read all three on 2026-10-03, so the honest answer is six confirmed by two Kyutai pages, plus Dutch listed by one.

No Korean, Japanese, Chinese or Hindi appears on any of those pages. If you need those, a self-hosted Pocket TTS is not the route today, and Sume's TTS language field is the one to look at.

What exactly does each Kyutai page list?

The pages were written at different times, which probably explains the mismatch: the blog post that announced the multilingual release is dated 2026-05-04, and the README may have moved since.

Pocket TTS language lists, read 2026-10-03 from Kyutai's GitHub README, Hugging Face model card and blog index.
LanguageREADMEModel cardBlog (2026-05-04)
EnglishListedListedListed
FrenchListedListedListed
GermanListedListedListed
PortugueseListedListedListed
ItalianListedListedListed
SpanishListedListedListed
DutchListedNot listedNot listed

How does Sume TTS pick a language?

Sume's TTS request has a language field: the language the voice speaks the transcript in, as a BCP-47 or ISO-639 code such as ko, ja or en. The schema says to set it for every non-English transcript, because an omitted value defaults to English at the provider. As a fallback Sume infers Korean or Japanese from a transcript that is only Hangul or kana, but the field is the reliable way.

There is also a guard rail: if the voice you picked does not match the language, the API returns a mismatch warning, and you resend with confirm_language_mismatch: true only after a person has agreed. The request never translates a transcript into English for you.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "transcript": "Votre commande est partie aujourd'"'"'hui.",
    "avatar_id": "'"$SUME_AVATAR_ID"'",
    "language": "fr",
    "mode": "sync",
    "wait_timeout_seconds": 30
  }'

What should you check before you rely on a language?

A language list says the model can speak it, not that every voice does it well. Before you ship a campaign in a language, run the same two sentences through the voice you plan to use and listen.

  • Generate a short sample in each target language with the real voice, not a demo voice.
  • For Sume, set language on every non-English request and keep the transcript in that language; the field does not translate.
  • For Pocket TTS, test Dutch separately, since only the README lists it.
  • Round-trip the output through speech-to-text and compare the words; Check a TTS take with an STT round trip shows the script.

How do you handle a mixed-language script?

A script that switches language mid-sentence is the case where language lists mislead most. A voice set to one language may read a foreign brand name in that language's sounds, which sometimes works and sometimes sounds wrong. The reliable approach is to split the script by language and synthesise each part with the right setting, then join the clips.

With Sume that is several TTS jobs, each with its own language, followed by one POST /v1/timeline-1.0/audio concat of up to 20 parts, which joins them in the sample domain with no re-synthesis and no silence at the seams. With a self-hosted Pocket TTS you would choose the language per call, using whichever of the six or seven languages the model supports for each part.

Plan the join in advance: keep every clip in the same channel layout (the concat refuses parts that differ), and render to wav if the joined file will feed a lip-sync or avatar step, because mp3 re-adds encoder padding at each edge. Join voiceover takes into one track has the full request. The README, the model card and the May blog post differ, and I could not tell from the text which is newest. Re-read the Kyutai blog before you plan around a list, because the project is moving; its most recent Pocket TTS post, dated 2026-09-28, was about the sampler, not languages.

Where does that leave a multilingual project?

For six European languages on your own CPU, Pocket TTS is a legitimate option. For a catalogue that includes Korean or Japanese, or when you want the audio file hosted and billed in one place, a Sume job with language set is the simpler path. Build an AI dubbing pipeline from STT, translate and TTS puts the language field to work end to end.

Sources

Related posts

More in Models

All Models posts

Written by Sume