Griffin clones a voice from about 10 seconds: what Sume does for voice

Tavus says Griffin can clone a voice from about 10 seconds of audio. Sume's TTS docs describe choosing a voice, not training one. Plan around that gap.

5 min readSume
All posts

Tavus says Griffin produces speech and can clone a voice from about 10 seconds of audio, but Griffin-Lite is open only to select trusted testers. In the Sume docs I read, TTS selects an existing voice (by avatar_handle, avatar_id or a voice id) and does not take a 10-second sample to train a new one. Plan your clip around a voice you can name.

What Tavus says

Tavus's Griffin page says speech generation produces audio packets as small as 10 ms, runs at a 48 kHz sample rate and can clone voices from roughly 10 seconds of audio. It also says Griffin-Lite is a research preview "not available for use for customers at this time", and that the same properties that make the model natural can let it deceive someone into thinking they are not talking to an AI. Tavus says it needs more alignment and safety work first.

So a voice cloned from 10 seconds is a capability to watch, not one you can order today.

What Sume's TTS takes

Sume's TTS surfaces take one text input, either transcript (1 to 20,000 characters) or a source-bound transcript_source, and one voice selector: a top-level avatar_id or avatar_handle, or a voice.id. language and output_format come with it. The TTS Router adds a required model from a Cartesia Sonic catalog (sonic-3.6 is the stable one that TTS 1.0 uses), and it bills characters at the list rate times 1.25.

The docs I read do not describe uploading a short sample to build a new voice, so I make no claim that Sume does.

Voice sources compared (read 2026-10-05)
QuestionGriffin-LiteSume TTS
Who can use itSelect trusted testers onlyPublic API
Voice originCloned from about 10 s of audioExisting voice chosen by avatar or voice id
Audio format48 kHz sample rateoutput_format field
Text limitNot stated on the page1 to 20,000 characters
Safety stanceMore disclosure features in developmentVoice is selected, not trained, in the docs I read

A safe plan today

Write the script, pick the voice you can name, generate the audio, then send it to a lip-sync route. Keep the audio inside the 5 to 14.8 second window for H3 Max lip sync, or use Fabric for other lengths.

Whatever voice you use, say that it is synthetic. See what to prepare for the Griffin-Lite safety hold for the consent and disclosure checklist.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume