Chatterbox Multilingual V3 voice cloning (MIT) vs Sume voice ids
Chatterbox is MIT-licensed TTS with 10-second voice cloning and Perth watermarking. How it differs from Sume, which takes a voice id and no clone upload.

Chatterbox clones a voice from a reference clip of about 10 seconds and runs on your own hardware, while Sume's public text-to-speech API has no clone-upload field at all: it takes a voice id that already exists in a Voices library. If you need to clone inside your own pipeline and keep the weights, Chatterbox is the closer fit; if you already have a voice id and want a hosted job, Sume is.
Chatterbox facts here come from the Resemble AI repository, read on 2026-10-02. Sume facts come from the Sume API reference.
What does Chatterbox ship today?
The repository describes an MIT-licensed family. Chatterbox Multilingual V3 (500M parameters) covers 23 or more languages, and the README lists codes such as ar, da, de, el, en, es, fi, fr, he, hi, it, ja and ko. Single-language models exist for Chinese, Spanish variants, Portuguese variants and Hindi.
Turbo (350M) and Nano (110M) models support native paralinguistic tags such as [cough], [laugh] and [chuckle]. The README says Nano runs on CPU at about three times faster than realtime on 8 cores, and recommends a standard GPU for Multilingual V3. The original model exposes exaggeration and cfg_weight to tune expressiveness. Every generated file carries Resemble's Perth neural watermark, which the README calls imperceptible.
How does voice cloning differ on Sume?
On Chatterbox you pass a reference clip at generation time. On Sume the request names a voice that already exists. The voice.id field accepts a TTS voice UUID or a voi_ library id, and the alternative is an avatar selector (avatar_id or avatar_handle) whose voice status is ready. The public API contract lists no route that creates a voice, so the clone step happens before you call it.
That has a practical consequence: a Sume request carries text and an id, not audio, so there is no reference clip to upload, store or lose. It also means you cannot A/B a new reference clip per request the way you can with a local model.
| Topic | Chatterbox | Sume TTS Router |
|---|---|---|
| Voice source | Reference clip of about 10 seconds at generation time | Existing voice id (UUID or voi_ id) or avatar voice |
| License | MIT code and models as listed in the repo | Hosted per-job service |
| Expressiveness | exaggeration and cfg_weight; paralinguistic tags on Turbo and Nano | generation_config with speed 0.6 to 1.5, volume 0.5 to 2 and an emotion guide |
| Watermark | Perth watermark on every file | Not specified in the API contract |
| Where it runs | Your GPU or CPU | Sume workers; you receive a media.sume.com URL |
What does a Sume call with an existing voice look like?
Submit with an idempotency key, then read the job. The default container is mp3 at 44.1 kHz and 128 kbps; ask for wav if the file will be edited again.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare two ways to get a voiceover.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en"
}'
Poll status and fetch the result:
curl https://api.sume.com/v1/jobs/$JOB_ID/status \
-H "Authorization: Bearer $SUME_API_KEY"
curl https://api.sume.com/v1/jobs/$JOB_ID/result \
-H "Authorization: Bearer $SUME_API_KEY"
Which should you choose?
Choose Chatterbox when voice cloning is the product: you want to experiment with reference clips, keep everything on your own machines, or need a watermark you control. Budget for hardware and for consent: cloning a voice you do not have permission for is a legal and platform-policy problem on any engine.
Choose Sume when the voice is already set and the work is the rest of the pipeline: billed jobs, durable URLs, word timestamps, sentence slices and joining takes with timeline audio. A Sume request also refuses a known language mismatch with 409 tts_voice_language_mismatch before any charge, which is a guardrail a local script does not give you by default.
What stays the same either way?
A short checklist applies to any cloned voice, local or hosted. Get written permission from the speaker for the specific uses you plan. Keep the reference recording and the permission together. Listen to a short test line in every language you will publish, because a clone that sounds right in English can drift in another language. Label synthetic speech where a platform or law asks you to. None of this depends on which engine renders the audio.
Both routes produce synthetic speech, so disclosure obligations do not depend on the engine. Keep a record of which voice, which model and which consent applies to each file. Sume's completed TTS job records the engine, the voice, the language and the settings it used, so you can reproduce a retake; with a local model, you have to log those yourself.
Sources
Related posts
More in Comparisons
- Cheapest 720p AI video per second: Veo, xAI, Sume
At 720p, Veo 3.1 lists $0.05 to $0.40 per second, xAI lists grok-imagine-video at $0.05, and Sume lists Grok Imagine Video 1.5 at $0.0125. Read 2026-10-01.
- Claude directory drops MCPB: remote vs local server, and Sume
Claude's directory no longer accepts MCPB desktop extensions. A remote HTTPS server needs no package; here is what that means for Sume.
- Claude Message Batches 100,000 requests vs a Sume bulk run of 100
A Claude Message Batch holds 100,000 requests or 256 MB. A Sume bulk run queues 1 to 100 Format runs with a concurrency window of 1 to 16. Plan accordingly.
- Claude batch canceled or expired results: billed? Sume cancel billing
Claude does not bill canceled or expired batch requests. Sume bills generation a run finished before a cancel, so a cancel stops spend but never refunds.
Written by Sume