Kokoro-82M (Apache 2.0) or a hosted TTS API: what Sume does instead
Kokoro-82M is an 82M-parameter Apache 2.0 TTS model with 54 voices in 8 languages. When to run it yourself and when Sume's hosted TTS route fits better.

Kokoro-82M is a small open text-to-speech model that you run on your own machine, while Sume's text-to-speech route is a hosted job you call over HTTPS and pay for per character. Pick Kokoro when you want free weights and full control of the runtime; pick a hosted route when you would rather not operate inference, voice storage and audio post-processing yourself.
The Kokoro facts below come from the model card on Hugging Face, read on 2026-10-02. The Sume facts come from the Sume API reference and the jobs and results guide.
What does the Kokoro-82M model card actually say?
The card lists an 82 million parameter model built on StyleTTS 2 with an ISTFTNet vocoder, released under the Apache 2.0 license. It supports 8 languages with 54 voices in total, outputs 24 kHz audio, and uses the misaki G2P library to turn text into phonemes. Voices are picked by name in the kokoro Python package, for example af_heart.
The card says training used a few hundred hours of audio from permissive or non-copyrighted sources and cost roughly $1,000. It does not describe voice cloning, a hosted API or an uptime commitment, so none of those should be assumed. Apache 2.0 allows commercial use of the weights, but you still own everything around them: the GPU or CPU, scaling, queueing and any audio clean-up.
How does that compare with Sume's TTS route?
Sume does not host Kokoro. Its public TTS surface is POST /v1/tts-router/generate, whose catalog today lists Cartesia Sonic engines only (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). You send text and a voice, get a job back and fetch a Sume-hosted audio file.
Read as a decision table:
| Question | Kokoro-82M (self-run) | Sume TTS Router (hosted) |
|---|---|---|
| Cost model | Your hardware and ops time | $0.0475 per 1,000 characters at the catalog price; spaces and punctuation count |
| Voices | 54 named voices in 8 languages | A voice id (UUID or voi_ library id) or an avatar's voice |
| Output | 24 kHz audio | mp3 44.1 kHz 128 kbps by default; wav and sample rates up to 48 kHz on request |
| Length | Whatever your code chunks | Up to 20,000 characters per request |
| License and rights | Apache 2.0 weights | Pay-per-job; your own rights checks still apply |
What does a hosted call look like?
A Sume job is asynchronous by default. Submit once with an Idempotency-Key, then read status and result from the shared jobs endpoints. The voice id must be a TTS voice UUID or a voi_ library id; any other shape is rejected with 400 invalid_voice_id before credits are reserved.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare two ways to get a voiceover.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en"
}'
Then poll the job:
curl https://api.sume.com/v1/jobs/$JOB_ID/status \
-H "Authorization: Bearer $SUME_API_KEY"
curl https://api.sume.com/v1/jobs/$JOB_ID/result \
-H "Authorization: Bearer $SUME_API_KEY"
When should you pick Kokoro, and when a hosted route?
Run Kokoro if you generate a lot of short English or multilingual lines offline, you already have a GPU or a fast CPU, and you do not need Sume's voice library or avatars. A permissive license and zero per-character cost are real advantages, and there is no network round trip.
Use a hosted route if you want billing per job, durable media.sume.com URLs, optional word timestamps and gapless sentence slices, and the same voice selector across TTS and avatar video. Say plainly what you give up: you cannot bring your own weights to Sume, and Sume's router does not carry Kokoro.
What should you check before shipping self-run audio?
Self-hosting moves the checklist onto you. None of the items below come from the model card; they are the questions to answer before a model becomes a production voice.
- Throughput: how many characters per minute your hardware renders, and what happens when ten requests arrive at once.
- Loudness and format: the card lists 24 kHz output, so decide whether you resample to 44.1 or 48 kHz before mixing with music or video.
- Pronunciation: names, brands and numbers need test lines. A phoneme library helps, but you still listen to every release.
- Disclosure: if the audio is synthetic speech in an ad or video, platforms and regulators may expect a label, whichever engine made it.
- Reproducibility: pin the package version and voice name, then store both with each render so a retake sounds like the first take.
Can you use both?
Yes. A common split is Kokoro for drafts and internal review, then a hosted voice for the final narration that goes into a video. Whatever produced the audio, you can join takes into one gapless file with timeline audio only when the files are already Sume-hosted, so import local files first. Keep the license of every engine in your records; the model card is the source of truth for Kokoro, and Sume's catalog is the source for its own price.
Sources
Related posts
More in Comparisons
- Leonardo.Ai API alternative for image generation: Sume Images
Leonardo's API submits to /generations and returns a generationId. Sume POST /v1/images works from a model catalog with idempotent retries. Compared.
- LinkedIn Brand Kit: single image ads only, colors, fonts and voice
LinkedIn's Brand Kit stores colors, heading and body fonts and a brand voice, for single image ads only. Here is what to prepare, and what Sume can do.
- Live vs file transcription cost: gpt-live-transcribe and batch rates
Streaming transcription costs 2 to 4 times file transcription on vendor pages. When live is worth it, and when Sume's $0.01 per minute file job is enough.
- LiveKit lists 16 avatar providers: real-time vs rendered Sume clips
LiveKit Agents lists 16 avatar providers that join a call as a participant. Sume Avatar 1.0 is not one of them; it renders finished clips. When to use which.
Written by Sume