Pocket TTS runs on 2 CPU cores: what ~200 ms first audio means
Kyutai's Pocket TTS lists 100M parameters, 2 CPU cores and ~200 ms to first audio. Whether that matters for a video voiceover, and a hosted TTS job's numbers.

Kyutai's Pocket TTS repository describes a 100-million-parameter text-to-speech model built to run on a CPU. Its page says it runs with only 2 CPU cores, returns the first audio chunk in about 200 ms, and runs about 6x faster than real time on a MacBook Air M4 CPU (read 2026-10-05). Those are the vendor's numbers on its hardware; your server will differ.
What the first-chunk number is for
First-chunk latency matters when someone is waiting for a reply: a voice assistant, a phone menu, a game character. For a 40-second video voiceover nobody hears the first 200 ms, because the whole file is mixed into a timeline before anyone plays it. What counts there is total time to a finished file, and whether the voice stays the same from one video to the next.
Two ways to get a voiceover file
| Question | Pocket TTS on your box | Sume TTS 1.0 job |
|---|---|---|
| Who runs it | You: Python 3.10 to 3.14, PyTorch 2.5+ | Sume, called with an API key |
| First audio | ~200 ms per the repo | Not streamed: you get a finished file from a job |
| Price | Your CPU time | $0.0475 per 1,000 characters, 1 cent minimum |
| Longest input | Your own limit | 20,000 characters per request |
| Output | What you write to disk | mp3, wav or raw, mp3 44.1 kHz 128 kbps by default |
What a hosted job looks like
Sume TTS is an async job: you submit text and a voice, then poll the job or take a webhook, and you get a file URL back. It is not a streaming endpoint, so a live reply loop is the wrong use. A 1,000-character narration (about a minute of speech) costs $0.0475, and a full 20,000-character request costs $0.95.
Budget it both ways
Take a channel with 30 videos a month and a 900-character script each. That is 27,000 characters, or about $1.28 of hosted TTS before each job rounds to a cent. A small CPU server you already pay for may be cheaper at that size; the trade is that you also own upgrades, voice files and failures at 2 a.m. At millions of characters a month the arithmetic shifts toward self-hosting, provided your script languages are supported.
A practical rule
Run Pocket TTS yourself if you already operate CPU servers and want audio generated next to your own data. Use a hosted job if the output goes into a video and you would rather not maintain a model runtime. Either way, test your own script: speed on a laptop says little about a shared VM.
- Check the repo's supported languages for your script before you commit; the page lists English, French, German, Portuguese, Italian, Spanish and Dutch, plus community-trained models for others.
- Keep the text you synthesize in version control so a re-render uses identical input.
Sources
Related posts
More in Models
- Qwen-Image 2.1 native RGBA vs Sume's ChatGPT Image 2.5 transparency
Qwen-Image 2.1's card lists native RGBA transparency under a research licence. On Sume, transparent output comes from ChatGPT Image 2.5's background field.
- Qwen-Image 2.1 is 7B: self-host it or call hosted qwen-image on Sume
Qwen-Image 2.1 is a 7B model under the Qwen Research License. Sume hosts qwen-image and qwen-image-max, not 2.1; hosted costs $0.025 or $0.094 per image.
- Recraft V4.1 from $0.035 vs V4 from $0.04: Sume lists only V4
Recraft's docs list V4.1 from $0.035 per image and V4 vectors from $0.04. Sume's catalog has recraft/recraft-v4 only, about $0.05 after its x 1.25 margin.
- 'reference_audio_urls requires a reference image or video'
Audio cannot be the only reference on Sume video models. Add one image or video reference. Which models take audio references, and their caps.
Written by Sume