Pocket TTS runs on 2 CPU cores: what ~200 ms first audio means

Kyutai's Pocket TTS lists 100M parameters, 2 CPU cores and ~200 ms to first audio. Whether that matters for a video voiceover, and a hosted TTS job's numbers.

4 min readSume
All posts

Kyutai's Pocket TTS repository describes a 100-million-parameter text-to-speech model built to run on a CPU. Its page says it runs with only 2 CPU cores, returns the first audio chunk in about 200 ms, and runs about 6x faster than real time on a MacBook Air M4 CPU (read 2026-10-05). Those are the vendor's numbers on its hardware; your server will differ.

What the first-chunk number is for

First-chunk latency matters when someone is waiting for a reply: a voice assistant, a phone menu, a game character. For a 40-second video voiceover nobody hears the first 200 ms, because the whole file is mixed into a timeline before anyone plays it. What counts there is total time to a finished file, and whether the voice stays the same from one video to the next.

Two ways to get a voiceover file

Self-hosted Pocket TTS vs a hosted Sume TTS job (read 2026-10-05)
QuestionPocket TTS on your boxSume TTS 1.0 job
Who runs itYou: Python 3.10 to 3.14, PyTorch 2.5+Sume, called with an API key
First audio~200 ms per the repoNot streamed: you get a finished file from a job
PriceYour CPU time$0.0475 per 1,000 characters, 1 cent minimum
Longest inputYour own limit20,000 characters per request
OutputWhat you write to diskmp3, wav or raw, mp3 44.1 kHz 128 kbps by default

What a hosted job looks like

Sume TTS is an async job: you submit text and a voice, then poll the job or take a webhook, and you get a file URL back. It is not a streaming endpoint, so a live reply loop is the wrong use. A 1,000-character narration (about a minute of speech) costs $0.0475, and a full 20,000-character request costs $0.95.

Budget it both ways

Take a channel with 30 videos a month and a 900-character script each. That is 27,000 characters, or about $1.28 of hosted TTS before each job rounds to a cent. A small CPU server you already pay for may be cheaper at that size; the trade is that you also own upgrades, voice files and failures at 2 a.m. At millions of characters a month the arithmetic shifts toward self-hosting, provided your script languages are supported.

A practical rule

Run Pocket TTS yourself if you already operate CPU servers and want audio generated next to your own data. Use a hosted job if the output goes into a video and you would rather not maintain a model runtime. Either way, test your own script: speed on a laptop says little about a shared VM.

  • Check the repo's supported languages for your script before you commit; the page lists English, French, German, Portuguese, Italian, Spanish and Dutch, plus community-trained models for others.
  • Keep the text you synthesize in version control so a re-render uses identical input.

Sources

Related posts

More in Models

All Models posts

Written by Sume