Pocket TTS license: the repo says MIT, not Apache-2.0
Some roundups call Kyutai's Pocket TTS Apache-2.0. Its GitHub page and LICENSE file read MIT-style. How to check a TTS license before you build on it.

A trend roundup this week listed Kyutai's Pocket TTS as Apache-2.0. The project's own GitHub page lists the license as MIT, and the first lines of its LICENSE file read "Permission is hereby granted, free of charge, to any person obtaining a copy of this software", which is the opening of the MIT license (both read 2026-10-05). When two sources disagree, the repository wins.
Why the difference is small but real
Both MIT and Apache-2.0 allow commercial use. They differ in details: Apache-2.0 includes an explicit patent grant and a notice requirement for changes, while MIT is shorter and has neither. If your legal team has a license allow-list, the label matters even when both are on it.
What the repo license does not tell you
- It covers the code in the repository. Model weights and voice prompts can carry their own terms, so read the model card where you download them.
- It says nothing about consent. Cloning a voice from an audio file is allowed by the software and still needs the speaker's permission.
- It does not cover the output you publish on a platform that requires AI-voice disclosure.
A five-minute license check for any TTS model
Open the repository's LICENSE file, not the README badge. Open the model card for the weights. Search both for words like non-commercial, research only and attribution. Write the date you checked in your project notes, because licenses change between releases.
How a wrong label travels
A roundup gets the license wrong, a blog copies it, and a procurement checklist now says Apache-2.0. Nobody lied, and nothing was verified. The fix costs one request: fetch the LICENSE file from the default branch, as this post did, and quote the first line. Re-check at each release, because maintainers can relicense future versions without changing the tag you pinned.
Where Sume fits
If you would rather not carry a model license at all, Sume TTS is a hosted API. You send a transcript and a voice id, pay $0.0475 per 1,000 characters, and the model terms stay Sume's concern. Voice consent is still yours: only synthesize voices you have the right to use.
Sources
Related posts
More in Models
- Pocket TTS runs on 2 CPU cores: what ~200 ms first audio means
Kyutai's Pocket TTS lists 100M parameters, 2 CPU cores and ~200 ms to first audio. Whether that matters for a video voiceover, and a hosted TTS job's numbers.
- Qwen-Image 2.1 native RGBA vs Sume's ChatGPT Image 2.5 transparency
Qwen-Image 2.1's card lists native RGBA transparency under a research licence. On Sume, transparent output comes from ChatGPT Image 2.5's background field.
- Qwen-Image 2.1 is 7B: self-host it or call hosted qwen-image on Sume
Qwen-Image 2.1 is a 7B model under the Qwen Research License. Sume hosts qwen-image and qwen-image-max, not 2.1; hosted costs $0.025 or $0.094 per image.
- Recraft V4.1 from $0.035 vs V4 from $0.04: Sume lists only V4
Recraft's docs list V4.1 from $0.035 per image and V4 vectors from $0.04. Sume's catalog has recraft/recraft-v4 only, about $0.05 after its x 1.25 margin.
Written by Sume