Pocket TTS speed: 6x on an M4 CPU, 2.3-2.5x on a VM CPU

Kyutai Pocket TTS runs about 6x real time on an M4 CPU and 2.3-2.5x on a 4-vCPU cloud VM. It cannot add silence for pauses. What to plan for.

5 min readSume
All posts

How fast is Kyutai Pocket TTS without a GPU? According to its GitHub README, about 6 times real time on a MacBook Air M4 CPU, with roughly 200 ms to the first audio chunk, using only 2 CPU cores. On a cloud x86 VM with a Tesla T4 the README gives about 2.3 to 2.5 times real time on CPU and about 6.28 times on the GPU.

The model has 100M parameters and an MIT license. The speed figures come from the project itself, so measure on your own hardware before you size a service around them.

What the README reports

The numbers depend on the machine, and the README is candid that a GPU does not always help. It notes that on hardware with strong single-thread CPU performance, such as Apple Silicon, it did not observe a GPU speedup, and that int8 dynamic quantization only works on CPU.

Pocket TTS figures from the project README (read 2026-10-03)
SetupReported speedNotes
MacBook Air M4, CPUAbout 6x real time2 CPU cores; about 200 ms to first chunk
Cloud x86 VM, CPUAbout 2.3-2.5x real time4 vCPUs, same instance as the T4 row
Same VM, Tesla T4 GPUAbout 6.28x real timeAbout 2.6x over that VM's CPU; no speedup seen on Apple Silicon

What a real-time factor means for a service

A speed of 2.3 times real time means one minute of audio takes under half a minute of compute on that machine. For a queue of narration jobs that is throughput, and you can estimate it with plain arithmetic: ten minutes of speech at 2.3 times real time is about 4.3 minutes of one VM's time, before start-up and any retries.

For a live agent the number that matters is first-chunk latency, which the README gives as about 200 ms on the M4. Add your speech recognition and language-model steps to that before you promise a response time.

The caveat that affects scripts

The README states that Pocket TTS does not support adding silence for pauses in text. If your script relies on a break tag or a long pause between lines, you have to build that in another way, for example by generating each line separately and joining them with silence yourself.

That join step is a common source of clicks, because compressed audio adds padding at each edge. If you stitch clips, use an uncompressed intermediate and a sample-accurate join.

A minimal call

The README's own Python example loads the model, picks one of its named voices and generates audio. It needs PyTorch 2.5 or later, and the install command is pip install pocket-tts.

from pocket_tts import TTSModel

tts_model = TTSModel.load_model()
voice_state = tts_model.get_state_for_audio_prompt("alba")
audio = tts_model.generate_audio(voice_state, "Hello world")

Self-host or call a job API

Running the model yourself buys control and no per-character fee, at the cost of owning capacity, updates and voice-consent handling. A hosted job API such as Sume's returns a result by polling or webhook, and its synchronous wait is capped at 30 seconds, as described in Sume jobs and results. A comparison of the two paths is in self-host or call a hosted API.

Sources

Related posts

More in Models

All Models posts

Written by Sume