Griffin clones a voice from about 10 seconds: what Sume does for voice
Tavus says Griffin can clone a voice from about 10 seconds of audio. Sume's TTS docs describe choosing a voice, not training one. Plan around that gap.

Tavus says Griffin produces speech and can clone a voice from about 10 seconds of audio, but Griffin-Lite is open only to select trusted testers. In the Sume docs I read, TTS selects an existing voice (by avatar_handle, avatar_id or a voice id) and does not take a 10-second sample to train a new one. Plan your clip around a voice you can name.
What Tavus says
Tavus's Griffin page says speech generation produces audio packets as small as 10 ms, runs at a 48 kHz sample rate and can clone voices from roughly 10 seconds of audio. It also says Griffin-Lite is a research preview "not available for use for customers at this time", and that the same properties that make the model natural can let it deceive someone into thinking they are not talking to an AI. Tavus says it needs more alignment and safety work first.
So a voice cloned from 10 seconds is a capability to watch, not one you can order today.
What Sume's TTS takes
Sume's TTS surfaces take one text input, either transcript (1 to 20,000 characters) or a source-bound transcript_source, and one voice selector: a top-level avatar_id or avatar_handle, or a voice.id. language and output_format come with it. The TTS Router adds a required model from a Cartesia Sonic catalog (sonic-3.6 is the stable one that TTS 1.0 uses), and it bills characters at the list rate times 1.25.
The docs I read do not describe uploading a short sample to build a new voice, so I make no claim that Sume does.
| Question | Griffin-Lite | Sume TTS |
|---|---|---|
| Who can use it | Select trusted testers only | Public API |
| Voice origin | Cloned from about 10 s of audio | Existing voice chosen by avatar or voice id |
| Audio format | 48 kHz sample rate | output_format field |
| Text limit | Not stated on the page | 1 to 20,000 characters |
| Safety stance | More disclosure features in development | Voice is selected, not trained, in the docs I read |
A safe plan today
Write the script, pick the voice you can name, generate the audio, then send it to a lip-sync route. Keep the audio inside the 5 to 14.8 second window for H3 Max lip sync, or use Fabric for other lengths.
Whatever voice you use, say that it is synthetic. See what to prepare for the Griffin-Lite safety hold for the consent and disclosure checklist.
Sources
Related posts
More in Comparisons
- Griffin-Lite or a rendered avatar clip for Q4? A decision table
Tavus Griffin-Lite is a research preview for invited testers. Sume Avatar 1.0 renders scripted clips from $11.04 per minute. Pick by the job, not the demo.
- MiniMax H3 vs H3 Max on Sume: 768p costs 33 percent more on Max
At 768p a second of minimax-h3-max costs $0.10 on Sume against $0.075 for minimax-h3. What the extra third buys, with a 10-second price table.
- Headliner Basic's 10 audiograms vs a podcast clip at $0.10 on Sume
Headliner Basic gives 10 unwatermarked audiograms a month for $9.99. A Sume Timeline clip is $0.10 per output minute. What you gain and what you give up.
- Avatar V needs 15 seconds of you; a Sume avatar needs one still
HeyGen Avatar V clones you from a 15-second webcam clip. Sume builds a reusable avatar from one photo for $0.95. What each needs and what you can render.
Written by Sume