Pocket TTS speed: 6x on an M4 CPU, 2.3-2.5x on a VM CPU
Kyutai Pocket TTS runs about 6x real time on an M4 CPU and 2.3-2.5x on a 4-vCPU cloud VM. It cannot add silence for pauses. What to plan for.

How fast is Kyutai Pocket TTS without a GPU? According to its GitHub README, about 6 times real time on a MacBook Air M4 CPU, with roughly 200 ms to the first audio chunk, using only 2 CPU cores. On a cloud x86 VM with a Tesla T4 the README gives about 2.3 to 2.5 times real time on CPU and about 6.28 times on the GPU.
The model has 100M parameters and an MIT license. The speed figures come from the project itself, so measure on your own hardware before you size a service around them.
What the README reports
The numbers depend on the machine, and the README is candid that a GPU does not always help. It notes that on hardware with strong single-thread CPU performance, such as Apple Silicon, it did not observe a GPU speedup, and that int8 dynamic quantization only works on CPU.
| Setup | Reported speed | Notes |
|---|---|---|
| MacBook Air M4, CPU | About 6x real time | 2 CPU cores; about 200 ms to first chunk |
| Cloud x86 VM, CPU | About 2.3-2.5x real time | 4 vCPUs, same instance as the T4 row |
| Same VM, Tesla T4 GPU | About 6.28x real time | About 2.6x over that VM's CPU; no speedup seen on Apple Silicon |
What a real-time factor means for a service
A speed of 2.3 times real time means one minute of audio takes under half a minute of compute on that machine. For a queue of narration jobs that is throughput, and you can estimate it with plain arithmetic: ten minutes of speech at 2.3 times real time is about 4.3 minutes of one VM's time, before start-up and any retries.
For a live agent the number that matters is first-chunk latency, which the README gives as about 200 ms on the M4. Add your speech recognition and language-model steps to that before you promise a response time.
The caveat that affects scripts
The README states that Pocket TTS does not support adding silence for pauses in text. If your script relies on a break tag or a long pause between lines, you have to build that in another way, for example by generating each line separately and joining them with silence yourself.
That join step is a common source of clicks, because compressed audio adds padding at each edge. If you stitch clips, use an uncompressed intermediate and a sample-accurate join.
A minimal call
The README's own Python example loads the model, picks one of its named voices and generates audio. It needs PyTorch 2.5 or later, and the install command is pip install pocket-tts.
from pocket_tts import TTSModel
tts_model = TTSModel.load_model()
voice_state = tts_model.get_state_for_audio_prompt("alba")
audio = tts_model.generate_audio(voice_state, "Hello world")Self-host or call a job API
Running the model yourself buys control and no per-character fee, at the cost of owning capacity, updates and voice-consent handling. A hosted job API such as Sume's returns a result by polling or webhook, and its synchronous wait is capped at 30 seconds, as described in Sume jobs and results. A comparison of the two paths is in self-host or call a hosted API.
Sources
Related posts
More in Models
- Polish, Dutch, Swedish, Turkish text to speech API: Sume pl nl sv tr
Sume's Voices library tags voices pl, nl, sv and tr alongside 12 other languages. What to send for each, and how Eleven v4's list compares.
- QuantFunc INT4 MiniMax H3: 3.2 s per step on an RTX 4090
QuantFunc's 4-bit MiniMax H3 claims 3.2 s per step on an RTX 4090. That is not a clip time. What the card says, what it omits, and when to use a hosted job.
- Qwen Image Edit Plus: $0.03 a megapixel for text edits
fal prices Qwen Image Edit Plus at $0.03 per megapixel and highlights text editing and multi-image input. What that means for a Sume edit workflow.
- Qwen Image Max is text-only on Sume; Qwen Image takes references
Qwen Image Max on Sume rejects reference images. For edits use qwen/qwen-image. The two ids compared: references, n, ratios, and the $0.075 Max list price.
Written by Sume