Chatterbox Multilingual V3 voice cloning (MIT) vs Sume voice ids

Chatterbox is MIT-licensed TTS with 10-second voice cloning and Perth watermarking. How it differs from Sume, which takes a voice id and no clone upload.

5 min readSume
All posts

Chatterbox clones a voice from a reference clip of about 10 seconds and runs on your own hardware, while Sume's public text-to-speech API has no clone-upload field at all: it takes a voice id that already exists in a Voices library. If you need to clone inside your own pipeline and keep the weights, Chatterbox is the closer fit; if you already have a voice id and want a hosted job, Sume is.

Chatterbox facts here come from the Resemble AI repository, read on 2026-10-02. Sume facts come from the Sume API reference.

What does Chatterbox ship today?

The repository describes an MIT-licensed family. Chatterbox Multilingual V3 (500M parameters) covers 23 or more languages, and the README lists codes such as ar, da, de, el, en, es, fi, fr, he, hi, it, ja and ko. Single-language models exist for Chinese, Spanish variants, Portuguese variants and Hindi.

Turbo (350M) and Nano (110M) models support native paralinguistic tags such as [cough], [laugh] and [chuckle]. The README says Nano runs on CPU at about three times faster than realtime on 8 cores, and recommends a standard GPU for Multilingual V3. The original model exposes exaggeration and cfg_weight to tune expressiveness. Every generated file carries Resemble's Perth neural watermark, which the README calls imperceptible.

How does voice cloning differ on Sume?

On Chatterbox you pass a reference clip at generation time. On Sume the request names a voice that already exists. The voice.id field accepts a TTS voice UUID or a voi_ library id, and the alternative is an avatar selector (avatar_id or avatar_handle) whose voice status is ready. The public API contract lists no route that creates a voice, so the clone step happens before you call it.

That has a practical consequence: a Sume request carries text and an id, not audio, so there is no reference clip to upload, store or lose. It also means you cannot A/B a new reference clip per request the way you can with a local model.

Chatterbox versus Sume TTS, read 2026-10-02
TopicChatterboxSume TTS Router
Voice sourceReference clip of about 10 seconds at generation timeExisting voice id (UUID or voi_ id) or avatar voice
LicenseMIT code and models as listed in the repoHosted per-job service
Expressivenessexaggeration and cfg_weight; paralinguistic tags on Turbo and Nanogeneration_config with speed 0.6 to 1.5, volume 0.5 to 2 and an emotion guide
WatermarkPerth watermark on every fileNot specified in the API contract
Where it runsYour GPU or CPUSume workers; you receive a media.sume.com URL

What does a Sume call with an existing voice look like?

Submit with an idempotency key, then read the job. The default container is mp3 at 44.1 kHz and 128 kbps; ask for wav if the file will be edited again.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: narration-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Welcome back. Today we compare two ways to get a voiceover.",
    "voice": { "id": "'"$VOICE_ID"'" },
    "language": "en"
  }'

Poll status and fetch the result:

curl https://api.sume.com/v1/jobs/$JOB_ID/status \
  -H "Authorization: Bearer $SUME_API_KEY"

curl https://api.sume.com/v1/jobs/$JOB_ID/result \
  -H "Authorization: Bearer $SUME_API_KEY"

Which should you choose?

Choose Chatterbox when voice cloning is the product: you want to experiment with reference clips, keep everything on your own machines, or need a watermark you control. Budget for hardware and for consent: cloning a voice you do not have permission for is a legal and platform-policy problem on any engine.

Choose Sume when the voice is already set and the work is the rest of the pipeline: billed jobs, durable URLs, word timestamps, sentence slices and joining takes with timeline audio. A Sume request also refuses a known language mismatch with 409 tts_voice_language_mismatch before any charge, which is a guardrail a local script does not give you by default.

What stays the same either way?

A short checklist applies to any cloned voice, local or hosted. Get written permission from the speaker for the specific uses you plan. Keep the reference recording and the permission together. Listen to a short test line in every language you will publish, because a clone that sounds right in English can drift in another language. Label synthetic speech where a platform or law asks you to. None of this depends on which engine renders the audio.

Both routes produce synthetic speech, so disclosure obligations do not depend on the engine. Keep a record of which voice, which model and which consent applies to each file. Sume's completed TTS job records the engine, the voice, the language and the settings it used, so you can reproduce a retake; with a local model, you have to log those yourself.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume