Kyutai Pocket TTS languages: what it speaks vs hosted Sume TTS

Pocket TTS is a 100M-parameter open model you run yourself. Its README lists seven languages; Sume's TTS is hosted, per character, with a voice id.

5 min readSume
All posts

Kyutai's Pocket TTS speaks English, French, German, Portuguese, Italian, Spanish and Dutch according to the README on GitHub, and you run it yourself on a CPU. Sume's TTS is the opposite deal: a hosted API you call with a transcript, a language code and a voice id, priced per character. For dubbing a video into several languages the question is whether you want to operate a model or call a job.

Everything about Pocket TTS below is from its GitHub README and its Hugging Face card, both read 2026-10-03. Everything about Sume is from the API reference.

What does Pocket TTS say it supports?

The README lists multi-language support for english, french, german, portuguese, italian, spanish and dutch, selected with a --language flag whose default is english. For the non-English languages it also mentions larger 24-layer variants that are higher quality but slower, chosen with a name such as italian_24l. The model is described as small, at 100M parameters.

The two pages disagree in places, which is worth knowing before you plan around either. The README says the license is MIT and names seven languages; the Hugging Face card, read the same day, names six languages (it omits Dutch) and shows a CC-BY-4.0 license, gated behind an agreement to usage conditions. If you ship a dub on it, read the page of the exact checkpoint you download.

Voice cloning is part of the tool: the README says the voice argument can take a plain wav file. It also warns that the quality of the sample is reproduced, so clean it first. The Hugging Face card lists voice impersonation or cloning without explicit lawful consent among its prohibited uses.

What does Sume's TTS give you instead?

Sume's POST /v1/tts-1.0/generate takes transcript up to 20,000 characters, an optional language code, and a voice by voice.id, avatar_id or avatar_handle. The request has no field for uploading a reference wav, so a clone is not made through that call. Results come back as jobs with word timings and optional per-sentence audio slices, which is what a dub needs to stay on its timing.

The schema publishes no language list and no count. It gives ko, ja and en as examples and says to set language for every non-English transcript. For each language you need, generate a test line and listen. Sume does not claim Dutch, or any of the seven, by name in its reference.

Side by side

Pocket TTS and Sume TTS 1.0, read 2026-10-03
QuestionPocket TTS (README)Sume TTS 1.0 (API reference)
Where it runsYour machine, CPUSume's hosted job API
LanguagesEnglish, French, German, Portuguese, Italian, Spanish, DutchNo published list; set language and test
VoiceBuilt-in voices or a wav you supplyA voice id or avatar; no wav upload field
Pauses from textSilence in the text input is not supportedWrite pauses through punctuation and sentence slices
Per-sentence timingsNot mentionedWord timings and sentence audio slices on request
CostYour compute$0.0475 per 1,000 characters (catalog rate)

When does running it yourself make sense?

Hosted wins when you want one request shape for every language, a job you can poll and retry with an idempotency key, and the sentence slices that feed a render. Those pieces are what turn speech into a finished dub: join the slices with Timeline audio, then render over the video with Timeline 1.0.

  • You dub in one of its seven languages, at volume, and already run CPU workers.
  • You need a cloned voice from your own sample and hold the consent for it.
  • Your pipeline is offline or you cannot send scripts to a hosted API.

Which licence and consent questions come with a dub?

Dubbing someone's video usually means a voice that is not theirs, and that is where licence and consent matter more than the language list. Pocket TTS ships with the ability to take a wav sample as the voice, so the consent question is yours to answer before you point it at anyone's recording. The Hugging Face card read 2026-10-03 lists impersonation without explicit lawful consent among prohibited uses.

On Sume's TTS the voice is chosen by id, so the question moves to who made that voice and for what. Use voices your workspace owns or has the right to use, and keep the reference for each language in your own records. Neither route removes the need to tell viewers when a voice is synthetic where a platform asks for that.

What to test before choosing

Run the same 20-second script through both in the language you ship. Compare the pronunciation of numbers and brand names, whether the speech fits the original time slot, and what a retry costs you in time. Pocket TTS shifts that cost to your hardware and your engineers; Sume shifts it to a per-character rate. Neither page offers a benchmark that settles quality for your script, and this post has none either.

Sources

Related posts

More in Models

All Models posts

Written by Sume