Pocket TTS voice cloning: a wav in, and what Sume does instead

Pocket TTS clones from a wav file you pass to --voice, with consent rules in its model card. Sume's API takes voice ids, not audio. Here is the difference.

5 min readSume
All posts

Pocket TTS clones a voice from a recording: the README says --voice accepts a plain wav file, and voices can also be exported as safetensors files and reused. Sume's public API does not take audio to clone. Its TTS routes take a voice id or an avatar whose voice is ready, and voices are made in the Sume app.

So if your plan is "upload a 10-second clip through an API and speak with that voice", Pocket TTS does it locally and Sume does not do it through the API. The rest of this post is about what each route needs from you, including the consent terms in Pocket TTS's model card.

How does Pocket TTS cloning work?

According to the GitHub README (read 2026-10-03), you point the --voice argument at a wav file, or at a safetensors export of a voice you made earlier. Nothing is uploaded anywhere unless you run it on a server you rent; the cloning happens on your CPU.

Pocket TTS cloning and use terms from the README and Hugging Face model card, read 2026-10-03.
QuestionWhat the page says
How do you clone?--voice takes a plain wav file; safetensors exports are also accepted
Is there a gate?The model card asks you to agree to usage conditions before access
Cloning someone without consent?Listed as a prohibited use: voice impersonation or cloning "without explicit lawful consent"
Deception?Prohibited: misinformation, disinformation, fraudulent calls, presenting generated content as genuine
LicenceThe README lists MIT; the model card lists CC-BY-4.0. Read both before shipping

Why does the licence line disagree between the two pages?

I read two Kyutai pages on 2026-10-03 and they do not say the same thing: the GitHub README lists MIT, and the Hugging Face model card lists CC-BY-4.0. They may cover different artifacts (code versus weights), but neither page said so in the text I read. I am not going to guess which applies to your use. Open both pages, and keep the attribution a CC-BY licence asks for if you ship the weights.

What does Sume take instead of a recording?

The Sume TTS request schema has two ways to pick a voice. One is avatar_id or avatar_handle: Sume resolves that avatar's voice when you submit, and the avatar's voice status must be ready. The other is voice.id, which must be a voice UUID or a Voices-library id starting with voi_; a voice name from another TTS product is rejected with a 400 before any credit is reserved. The OpenAPI document has no endpoint that accepts audio to build a voice.

The practical route is to create the voice once in the Sume app, then call the API with the avatar or the id for every clip after that. Clone your voice once and reuse it in TTS walks through the app side, and AI voice from a text description covers a voice with no recording at all.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "transcript": "This line is spoken by the voice I made in the app.",
    "avatar_handle": "'"$SUME_AVATAR_HANDLE"'",
    "mode": "sync",
    "wait_timeout_seconds": 30
  }'

What belongs in a consent record?

Pocket TTS's model card ties cloning to explicit lawful consent, and the same rule is sensible for any pipeline. A useful record is short: the speaker's name, the date, the specific uses they agreed to (for example product videos but not phone calls), how long the voice may be used, and how they can withdraw. Keep the original reference recording with the record so you can show which clip a voice came from.

Store the consent next to the voice, not in a separate folder nobody opens. If a voice lives in the Sume app, note its avatar handle or voice id in the record; if it lives as a safetensors file on a server, note the file name and the machine. When the speaker withdraws, you know exactly what to delete.

Cloned voices also raise disclosure questions when the audio reaches an audience. Two related posts cover the legal and platform side: Cartesia's acceptable-use consent clause and the best voice cloning TTS choices in 2026.

Which should you use?

The decision is about where the recording is allowed to live and who has to approve it.

  • Choose Pocket TTS when the reference clip must never leave your machine, or when you want to experiment with many reference voices quickly and you own the consent paperwork.
  • Choose a Sume voice when the speaker is a person or brand you have recorded once, and every later clip should reuse the same voice from any script, in any step of a Sume pipeline.
  • In both cases keep a written consent record from the speaker. Pocket TTS's own card calls out cloning without explicit lawful consent as prohibited, and that is a good floor for any voice you make.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume