Why TTS plus lip-sync avatars feel fake: Griffin's 2.4% vs 48%
Tavus says its older avatar stack passed as human 2.4% of the time and Griffin-Lite 48%. Here is what that means for TTS-then-lip-sync clips.
Tavus's study reports that its previous stack, Phoenix-4.5 with Sparrow-2 and Raven-1, passed as human 2.4% of the time, while Griffin-Lite passed 48% (26 of 54 participants) after one-minute calls. The jump is credited to one model that perceives, decides and generates together, instead of separate parts handing off. A Sume clip is a pipeline too, so the fixes you control are timing, framing and honesty.
What the numbers say
On the Griffin page, Tavus describes Griffin as a single system that combines perception, conversation and expressive video generation. It contrasts that with the earlier three-model setup and reports the 2.4% and 48% pass rates. Participants who said "real" were 79% confident and those who said "AI" were 81% confident.
The sample is 54 people in short calls, which is small, so treat 48% as a signal about direction, not a measured rate for your audience.
Where a TTS-then-lip-sync clip gives itself away
Sume's guidance for talking shots is TTS first, then a lip-sync model, never a video model with narration laid over it. That keeps the mouth tied to the audio, but the handoff leaves seams you can see.
Three are common: even pacing with no pauses, a face that never reacts while it is silent, and clip boundaries where the posture resets. You can reduce each one without a new model.
- Pacing: write shorter sentences and vary their length, then add
voice.type: "silence"beats in multi-scenevideo_inputs, wheredurationis required and no script is allowed. - One lip-sync model per run, so frame rate and skin tone do not change between clips.
- Check sync with word times: compare transcript timestamps against frames.
- State that the video is synthetic, in a caption or at the start.
| Seam | What a viewer notices | Fix on Sume |
|---|---|---|
| Even pacing | Reads like a recording of a script | Silence beats, shorter sentences |
| Frozen face while idle | No reaction to anything | Scene prompt, short clips that cut away |
| Model switch mid-video | Look and fps jump | One lip-sync model per run |
| Mouth drift | Words and lips disagree | Transcript word times vs frames |
The honest conclusion
Do not chase a pass rate. Make the clip good, and label it. See how to read the Griffin 48% study and silence beats in a recorded avatar video.
Sources
Related posts
More in Use cases
- Win-back video for lapsed customers: 30-second avatar script, cost
A win-back email clip made with a Sume avatar: what to say in 30 seconds, what not to promise, and the cost per lapsed account on standard, plus and max.
- X carousel ads: 2 to 6 media assets, rendered as one Sume batch
X carousel cards hold 2 to 6 media assets. Make the whole set at one aspect ratio in a single Sume Image API call with n, then check each file.
- X image ad specs: 1:1 or 1.91:1, 800px wide, 3 MB, from Sume
X website image cards take 1:1 or 1.91:1, at least 800px wide and 3 MB. Render 2:1 on Sume, crop to 1.91:1 with Pillow, and check size before upload.
- X lead generation form: name up to 255 characters, purpose 1000
X's Ads API lead generation docs cap a form name at 1 to 255 characters and its purpose at 1 to 1000. Validate copy variants before you create the form.
Written by Sume