Why TTS plus lip-sync avatars feel fake: Griffin's 2.4% vs 48%

Tavus says its older avatar stack passed as human 2.4% of the time and Griffin-Lite 48%. Here is what that means for TTS-then-lip-sync clips.

5 min readSume
All posts

Tavus's study reports that its previous stack, Phoenix-4.5 with Sparrow-2 and Raven-1, passed as human 2.4% of the time, while Griffin-Lite passed 48% (26 of 54 participants) after one-minute calls. The jump is credited to one model that perceives, decides and generates together, instead of separate parts handing off. A Sume clip is a pipeline too, so the fixes you control are timing, framing and honesty.

What the numbers say

On the Griffin page, Tavus describes Griffin as a single system that combines perception, conversation and expressive video generation. It contrasts that with the earlier three-model setup and reports the 2.4% and 48% pass rates. Participants who said "real" were 79% confident and those who said "AI" were 81% confident.

The sample is 54 people in short calls, which is small, so treat 48% as a signal about direction, not a measured rate for your audience.

Where a TTS-then-lip-sync clip gives itself away

Sume's guidance for talking shots is TTS first, then a lip-sync model, never a video model with narration laid over it. That keeps the mouth tied to the audio, but the handoff leaves seams you can see.

Three are common: even pacing with no pauses, a face that never reacts while it is silent, and clip boundaries where the posture resets. You can reduce each one without a new model.

  • Pacing: write shorter sentences and vary their length, then add voice.type: "silence" beats in multi-scene video_inputs, where duration is required and no script is allowed.
  • One lip-sync model per run, so frame rate and skin tone do not change between clips.
  • Check sync with word times: compare transcript timestamps against frames.
  • State that the video is synthetic, in a caption or at the start.
Pipeline seams and the fix you control (read 2026-10-05)
SeamWhat a viewer noticesFix on Sume
Even pacingReads like a recording of a scriptSilence beats, shorter sentences
Frozen face while idleNo reaction to anythingScene prompt, short clips that cut away
Model switch mid-videoLook and fps jumpOne lip-sync model per run
Mouth driftWords and lips disagreeTranscript word times vs frames

The honest conclusion

Do not chase a pass rate. Make the clip good, and label it. See how to read the Griffin 48% study and silence beats in a recorded avatar video.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume