Kyutai Pocket TTS languages: what it speaks vs hosted Sume TTS
Pocket TTS is a 100M-parameter open model you run yourself. Its README lists seven languages; Sume's TTS is hosted, per character, with a voice id.

Kyutai's Pocket TTS speaks English, French, German, Portuguese, Italian, Spanish and Dutch according to the README on GitHub, and you run it yourself on a CPU. Sume's TTS is the opposite deal: a hosted API you call with a transcript, a language code and a voice id, priced per character. For dubbing a video into several languages the question is whether you want to operate a model or call a job.
Everything about Pocket TTS below is from its GitHub README and its Hugging Face card, both read 2026-10-03. Everything about Sume is from the API reference.
What does Pocket TTS say it supports?
The README lists multi-language support for english, french, german, portuguese, italian, spanish and dutch, selected with a --language flag whose default is english. For the non-English languages it also mentions larger 24-layer variants that are higher quality but slower, chosen with a name such as italian_24l. The model is described as small, at 100M parameters.
The two pages disagree in places, which is worth knowing before you plan around either. The README says the license is MIT and names seven languages; the Hugging Face card, read the same day, names six languages (it omits Dutch) and shows a CC-BY-4.0 license, gated behind an agreement to usage conditions. If you ship a dub on it, read the page of the exact checkpoint you download.
Voice cloning is part of the tool: the README says the voice argument can take a plain wav file. It also warns that the quality of the sample is reproduced, so clean it first. The Hugging Face card lists voice impersonation or cloning without explicit lawful consent among its prohibited uses.
What does Sume's TTS give you instead?
Sume's POST /v1/tts-1.0/generate takes transcript up to 20,000 characters, an optional language code, and a voice by voice.id, avatar_id or avatar_handle. The request has no field for uploading a reference wav, so a clone is not made through that call. Results come back as jobs with word timings and optional per-sentence audio slices, which is what a dub needs to stay on its timing.
The schema publishes no language list and no count. It gives ko, ja and en as examples and says to set language for every non-English transcript. For each language you need, generate a test line and listen. Sume does not claim Dutch, or any of the seven, by name in its reference.
Side by side
| Question | Pocket TTS (README) | Sume TTS 1.0 (API reference) |
|---|---|---|
| Where it runs | Your machine, CPU | Sume's hosted job API |
| Languages | English, French, German, Portuguese, Italian, Spanish, Dutch | No published list; set language and test |
| Voice | Built-in voices or a wav you supply | A voice id or avatar; no wav upload field |
| Pauses from text | Silence in the text input is not supported | Write pauses through punctuation and sentence slices |
| Per-sentence timings | Not mentioned | Word timings and sentence audio slices on request |
| Cost | Your compute | $0.0475 per 1,000 characters (catalog rate) |
When does running it yourself make sense?
Hosted wins when you want one request shape for every language, a job you can poll and retry with an idempotency key, and the sentence slices that feed a render. Those pieces are what turn speech into a finished dub: join the slices with Timeline audio, then render over the video with Timeline 1.0.
- You dub in one of its seven languages, at volume, and already run CPU workers.
- You need a cloned voice from your own sample and hold the consent for it.
- Your pipeline is offline or you cannot send scripts to a hosted API.
Which licence and consent questions come with a dub?
Dubbing someone's video usually means a voice that is not theirs, and that is where licence and consent matter more than the language list. Pocket TTS ships with the ability to take a wav sample as the voice, so the consent question is yours to answer before you point it at anyone's recording. The Hugging Face card read 2026-10-03 lists impersonation without explicit lawful consent among prohibited uses.
On Sume's TTS the voice is chosen by id, so the question moves to who made that voice and for what. Use voices your workspace owns or has the right to use, and keep the reference for each language in your own records. Neither route removes the need to tell viewers when a voice is synthetic where a platform asks for that.
What to test before choosing
Run the same 20-second script through both in the language you ship. Compare the pronunciation of numbers and brand names, whether the speech fits the original time slot, and what a retry costs you in time. Pocket TTS shifts that cost to your hardware and your engineers; Sume shifts it to a per-character rate. Neither page offers a benchmark that settles quality for your script, and this post has none either.
Sources
Related posts
More in Models
- Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.
- Omni Flash vs H3 Max for reference-to-video: limits side by side
Gemini Omni Flash 1.1 takes 10 images and 3 short videos; MiniMax H3 Max adds audio refs, 12 files total. Limits, tags and a sample request on Sume.
- Shortest AI video clip via API: 2 seconds on Wan 3.0 only
Sume's video catalog by minimum duration: Wan 3.0 takes 2 seconds, Omni 3, Seedance and Kling 4, MiniMax H3 5. Plus how to trim below a floor.
- Silent AI video: generate_audio false or drop the audio after
Seedance, Kling and Wan take generate_audio false; Omni and MiniMax H3 always make sound. How to get a silent clip on Sume, and what it costs.
Written by Sume