Pocket TTS voice cloning: a wav in, and what Sume does instead
Pocket TTS clones from a wav file you pass to --voice, with consent rules in its model card. Sume's API takes voice ids, not audio. Here is the difference.

Pocket TTS clones a voice from a recording: the README says --voice accepts a plain wav file, and voices can also be exported as safetensors files and reused. Sume's public API does not take audio to clone. Its TTS routes take a voice id or an avatar whose voice is ready, and voices are made in the Sume app.
So if your plan is "upload a 10-second clip through an API and speak with that voice", Pocket TTS does it locally and Sume does not do it through the API. The rest of this post is about what each route needs from you, including the consent terms in Pocket TTS's model card.
How does Pocket TTS cloning work?
According to the GitHub README (read 2026-10-03), you point the --voice argument at a wav file, or at a safetensors export of a voice you made earlier. Nothing is uploaded anywhere unless you run it on a server you rent; the cloning happens on your CPU.
| Question | What the page says |
|---|---|
| How do you clone? | --voice takes a plain wav file; safetensors exports are also accepted |
| Is there a gate? | The model card asks you to agree to usage conditions before access |
| Cloning someone without consent? | Listed as a prohibited use: voice impersonation or cloning "without explicit lawful consent" |
| Deception? | Prohibited: misinformation, disinformation, fraudulent calls, presenting generated content as genuine |
| Licence | The README lists MIT; the model card lists CC-BY-4.0. Read both before shipping |
Why does the licence line disagree between the two pages?
I read two Kyutai pages on 2026-10-03 and they do not say the same thing: the GitHub README lists MIT, and the Hugging Face model card lists CC-BY-4.0. They may cover different artifacts (code versus weights), but neither page said so in the text I read. I am not going to guess which applies to your use. Open both pages, and keep the attribution a CC-BY licence asks for if you ship the weights.
What does Sume take instead of a recording?
The Sume TTS request schema has two ways to pick a voice. One is avatar_id or avatar_handle: Sume resolves that avatar's voice when you submit, and the avatar's voice status must be ready. The other is voice.id, which must be a voice UUID or a Voices-library id starting with voi_; a voice name from another TTS product is rejected with a 400 before any credit is reserved. The OpenAPI document has no endpoint that accepts audio to build a voice.
The practical route is to create the voice once in the Sume app, then call the API with the avatar or the id for every clip after that. Clone your voice once and reuse it in TTS walks through the app side, and AI voice from a text description covers a voice with no recording at all.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "This line is spoken by the voice I made in the app.",
"avatar_handle": "'"$SUME_AVATAR_HANDLE"'",
"mode": "sync",
"wait_timeout_seconds": 30
}'
What belongs in a consent record?
Pocket TTS's model card ties cloning to explicit lawful consent, and the same rule is sensible for any pipeline. A useful record is short: the speaker's name, the date, the specific uses they agreed to (for example product videos but not phone calls), how long the voice may be used, and how they can withdraw. Keep the original reference recording with the record so you can show which clip a voice came from.
Store the consent next to the voice, not in a separate folder nobody opens. If a voice lives in the Sume app, note its avatar handle or voice id in the record; if it lives as a safetensors file on a server, note the file name and the machine. When the speaker withdraws, you know exactly what to delete.
Cloned voices also raise disclosure questions when the audio reaches an audience. Two related posts cover the legal and platform side: Cartesia's acceptable-use consent clause and the best voice cloning TTS choices in 2026.
Which should you use?
The decision is about where the recording is allowed to live and who has to approve it.
- Choose Pocket TTS when the reference clip must never leave your machine, or when you want to experiment with many reference voices quickly and you own the consent paperwork.
- Choose a Sume voice when the speaker is a person or brand you have recorded once, and every later clip should reuse the same voice from any script, in any step of a Sume pipeline.
- In both cases keep a written consent record from the speaker. Pocket TTS's own card calls out cloning without explicit lawful consent as prohibited, and that is a good floor for any voice you make.
Sources
Related posts
More in Comparisons
- Schedule a weekly AI video: Sume Scheduled, Hermes cron or API
Three ways to run an AI video on a weekly clock with Sume: a dashboard schedule, a Hermes cron job, or plain cron calling a Format. What each can and cannot do.
- Which image and video models on Sume have open weights?
Sume's catalog is hosted. Where vendors publish weights for families Sume lists (Qwen-Image, MiniMax H3, FLUX.2) and why a null hugging_face_id proves nothing.
- Wix Stores promo video maker vs a custom AI product clip
Wix Stores can auto-make a promo video per product, free. When is a custom Sume image-to-video clip from your own photo worth the extra step? A side-by-side.
- YouTube reused vs inauthentic content: which rule hits AI Shorts?
Two YouTube monetization rules and one new Shorts reach change, side by side. Which one an AI-generated Short can trip, and what to vary to avoid each.
Written by Sume