Pocket TTS API: run it yourself or call a hosted TTS API
Kyutai's Pocket TTS installs with pip and serves from localhost. If you want a hosted API with job URLs instead, here is the Sume request and what changes.

Pocket TTS does not come as a hosted API from Kyutai. Its README documents a local install (pip install pocket-tts), a pocket-tts generate command, and a pocket-tts serve command that opens a web interface on localhost:8000. So "the Pocket TTS API" means a server you run yourself, on a CPU machine you pay for.
If what you want is a text-to-speech API you call with a key and get a file back from, Sume's POST /v1/tts-1.0/generate is that: a job you submit, poll or receive by webhook, and a hosted audio URL at the end. This post lines the two up so you can pick on facts, not on the word "API".
What does the Pocket TTS README promise?
Kyutai's Pocket TTS launched on 2026-01-13 as a small model that runs on CPU, and its repository describes it in these terms. All values below are the README's own, read 2026-10-03; the speed figure is Kyutai's number on one laptop, not a measurement of yours.
| Item | README or blog says |
|---|---|
| Size | 100M parameters |
| Hardware | CPU, 2 cores, no GPU needed |
| Speed | About 200 ms to the first audio chunk; about 6x real time on a MacBook Air M4 CPU |
| Run it | pocket-tts generate for a file, pocket-tts serve for a local web interface at localhost:8000, or TTSModel and generate_audio() in Python |
| Streaming | Supported, per the README |
| Input length | "Infinitely long text inputs" |
| Cloning | Pass a wav file to --voice |
What does a hosted job look like instead?
Sume TTS 1.0 is not streaming. The OpenAPI description says the route is an async job with polling or a webhook, and that completed results expose mirrored audio artifacts. mode: sync waits up to 30 seconds for the finished job and otherwise hands you the polling URLs; that wait bounds the HTTP call, not the job.
The request needs a transcript (up to 20,000 characters) and a voice. A voice comes from avatar_id or avatar_handle (list GET /v1/avatar-1.0/avatars and use one whose voice status is ready) or from a voice.id. The default output is mp3 at 44.1 kHz and 128 kbps. Here is the whole round trip in Python; set SUME_API_KEY and SUME_AVATAR_ID first.
import os, time, requests
BASE, H = "https://api.sume.com", {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
h = dict(H, **({"Idempotency-Key": key} if key else {}))
d = requests.post(BASE + path, headers=h, json=body, timeout=60)
d.raise_for_status()
d = d.json()["data"]
while not d["terminal"]:
time.sleep(d.get("next_poll_after_seconds") or 2)
d = requests.get(d["status_url"], headers=H, timeout=30).json()["data"]
r = requests.get(d["result_url"], headers=H, timeout=30)
r.raise_for_status() # failed or canceled jobs answer 409 here
return r.json()["data"]["result"]
result = run("/v1/tts-1.0/generate", {
"transcript": "Your order shipped today and arrives on Friday.",
"avatar_id": os.environ["SUME_AVATAR_ID"],
"mode": "async",
})
print(result["artifacts"][0]["url"])
What do you give up on each side?
Self-hosting Pocket TTS buys control: your hardware, your text never leaves the box, and the latency to a first chunk is whatever your CPU gives you. It also hands you the operations: a process to keep up, scaling when two users talk at once, and updates when Kyutai ships a new language or sampler. Kyutai's recent posts show how fast it moves: six languages in May, training code in August, a new sampler head on 2026-09-28.
- Pick Pocket TTS when audio must be generated on your own machine, you need speech while the text is still being written, or you want to train or fine-tune with Kyutai's released training code.
- Pick a hosted job when you want a durable audio URL, one bill, and a file you can hand to the next step (Sume's timeline, avatar and caption jobs take Sume-hosted audio).
- Do not pick a hosted job for a live conversation. Sume returns a finished file, so the first sound arrives after the whole clip is synthesised; the streaming TTS post shows how to cut a script into sentence jobs to start playback sooner.
How should you test both before choosing?
Take ten real lines from your product, including a long one, a number-heavy one and one in a second language. Run them through pocket-tts generate on the machine you would deploy to, and through one Sume job each. Listen blind, in a shuffled order, with the person who will own the voice quality.
Then time the part you care about. For Pocket TTS that is the first chunk on your own CPU under load, not the README's figure from one laptop. For Sume it is the time from submit to a ready file, which depends on queue state and clip length, so measure a few dozen jobs rather than one. A fixed 30-second sync wait is an upper bound on one HTTP call, not an expected duration.
Finally write down who is on call when the voice stops working. With a self-hosted model that is your team; with a hosted job it is a status code, a job id to log, and a retry. The realtime voice or async TTS comparison works through the same trade for narration.
What does the hosted side cost?
Sume bills TTS by character: Cartesia's list rate in the pricing code is 38 micro-dollars per character, times Sume's 1.25 margin, rounded up to whole cents per job. That is about $0.0475 per 1,000 characters before rounding, so a 1,000-character job is $0.05 and a 20,000-character job is $0.95. GET /v1/tts-router/models returns the list rate and billing formula live, so read it before you budget; the price-per-1,000-characters comparison puts it next to other vendors.
Pocket TTS has no per-character price, only the cost of the machine, which is why it wins for very high volume on hardware you already own and loses for a few clips a day.
Sources
Related posts
More in Developers
- Prefect 3 task retries for a Sume job: same Idempotency-Key
Retry a Sume image job in Prefect 3 without paying twice: a tested flow with retry_condition_fn, delay list and an order-derived idempotency key.
- Pydantic model to Sume output_schema: extra forbid, no defaults
Turn a Pydantic v2 model into a valid Sume Format output_schema: extra=forbid, nullable instead of defaults, and the SumeMediaFile reference. Tested.
- Unit test Sume webhook signature checks in pytest (Python)
A pytest file for HMAC webhook verification: tamper, rotation header, stale timestamp and empty secret, written against Sume's sume-v1 scheme.
- Log x-sume-request-id and Idempotency-Key on every call (Python)
A requests response hook that writes one JSON log line per Sume call: x-sume-request-id, idempotency key, error code and rate-limit headers. Tested.
Written by Sume