Kokoro-82M (Apache 2.0) or a hosted TTS API: what Sume does instead

Kokoro-82M is an 82M-parameter Apache 2.0 TTS model with 54 voices in 8 languages. When to run it yourself and when Sume's hosted TTS route fits better.

5 min readSume
All posts

Kokoro-82M is a small open text-to-speech model that you run on your own machine, while Sume's text-to-speech route is a hosted job you call over HTTPS and pay for per character. Pick Kokoro when you want free weights and full control of the runtime; pick a hosted route when you would rather not operate inference, voice storage and audio post-processing yourself.

The Kokoro facts below come from the model card on Hugging Face, read on 2026-10-02. The Sume facts come from the Sume API reference and the jobs and results guide.

What does the Kokoro-82M model card actually say?

The card lists an 82 million parameter model built on StyleTTS 2 with an ISTFTNet vocoder, released under the Apache 2.0 license. It supports 8 languages with 54 voices in total, outputs 24 kHz audio, and uses the misaki G2P library to turn text into phonemes. Voices are picked by name in the kokoro Python package, for example af_heart.

The card says training used a few hundred hours of audio from permissive or non-copyrighted sources and cost roughly $1,000. It does not describe voice cloning, a hosted API or an uptime commitment, so none of those should be assumed. Apache 2.0 allows commercial use of the weights, but you still own everything around them: the GPU or CPU, scaling, queueing and any audio clean-up.

How does that compare with Sume's TTS route?

Sume does not host Kokoro. Its public TTS surface is POST /v1/tts-router/generate, whose catalog today lists Cartesia Sonic engines only (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). You send text and a voice, get a job back and fetch a Sume-hosted audio file.

Read as a decision table:

Kokoro-82M versus Sume TTS Router, read 2026-10-02
QuestionKokoro-82M (self-run)Sume TTS Router (hosted)
Cost modelYour hardware and ops time$0.0475 per 1,000 characters at the catalog price; spaces and punctuation count
Voices54 named voices in 8 languagesA voice id (UUID or voi_ library id) or an avatar's voice
Output24 kHz audiomp3 44.1 kHz 128 kbps by default; wav and sample rates up to 48 kHz on request
LengthWhatever your code chunksUp to 20,000 characters per request
License and rightsApache 2.0 weightsPay-per-job; your own rights checks still apply

What does a hosted call look like?

A Sume job is asynchronous by default. Submit once with an Idempotency-Key, then read status and result from the shared jobs endpoints. The voice id must be a TTS voice UUID or a voi_ library id; any other shape is rejected with 400 invalid_voice_id before credits are reserved.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: narration-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Welcome back. Today we compare two ways to get a voiceover.",
    "voice": { "id": "'"$VOICE_ID"'" },
    "language": "en"
  }'

Then poll the job:

curl https://api.sume.com/v1/jobs/$JOB_ID/status \
  -H "Authorization: Bearer $SUME_API_KEY"

curl https://api.sume.com/v1/jobs/$JOB_ID/result \
  -H "Authorization: Bearer $SUME_API_KEY"

When should you pick Kokoro, and when a hosted route?

Run Kokoro if you generate a lot of short English or multilingual lines offline, you already have a GPU or a fast CPU, and you do not need Sume's voice library or avatars. A permissive license and zero per-character cost are real advantages, and there is no network round trip.

Use a hosted route if you want billing per job, durable media.sume.com URLs, optional word timestamps and gapless sentence slices, and the same voice selector across TTS and avatar video. Say plainly what you give up: you cannot bring your own weights to Sume, and Sume's router does not carry Kokoro.

What should you check before shipping self-run audio?

Self-hosting moves the checklist onto you. None of the items below come from the model card; they are the questions to answer before a model becomes a production voice.

  • Throughput: how many characters per minute your hardware renders, and what happens when ten requests arrive at once.
  • Loudness and format: the card lists 24 kHz output, so decide whether you resample to 44.1 or 48 kHz before mixing with music or video.
  • Pronunciation: names, brands and numbers need test lines. A phoneme library helps, but you still listen to every release.
  • Disclosure: if the audio is synthetic speech in an ad or video, platforms and regulators may expect a label, whichever engine made it.
  • Reproducibility: pin the package version and voice name, then store both with each render so a retake sounds like the first take.

Can you use both?

Yes. A common split is Kokoro for drafts and internal review, then a hosted voice for the final narration that goes into a video. Whatever produced the audio, you can join takes into one gapless file with timeline audio only when the files are already Sume-hosted, so import local files first. Keep the license of every engine in your records; the model card is the source of truth for Kokoro, and Sume's catalog is the source for its own price.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume