AI video Turing test: how to read Griffin's 48% study

Tavus says 48% of 54 people took Griffin for a real person after a one-minute call. What the study shows, and how to disclose AI in clips you render.

5 min readSume
All posts

Tavus reports that 48% of people, 26 of 54, believed Griffin was a real person after a one-minute video call, against 2.4% (1 of 41) for previous systems (Tavus Griffin post, read 2026-10-03). It is a vendor-run study with a short call and a small group, so read it as a signal about conversational behavior, not as proof that a video call can no longer be trusted.

The more useful lesson for anyone shipping avatar video, live or rendered, is that audiences may not be able to tell, so disclosure has to come from you.

What did the Tavus study actually measure?

The numbers in the Tavus post are narrow. The Turing-style result is one-minute calls. The separate benchmark result is a 3.83 generation score on NVIDIA's VideoFDB, against a human reference of 3.92 (press release, read 2026-10-03). The Tavus post also reports a perception score of 3.73 and 0.43 seconds of video latency on H100s.

A third-party analysis that we cite only as reported notes the protocol is vendor-conducted and says the evaluations cover conversational realism, not accurate advice, reliable account actions or prompt-injection resistance (kingy.ai, read 2026-10-03). That is the right framing: passing for human for a minute says nothing about whether the agent gave a correct answer.

Reported Griffin figures and what each does not tell you, read 2026-10-03
Reported figureSourceWhat it does not tell you
48% believed it was a real person (26 of 54)Tavus Griffin postAnything about calls longer than one minute
2.4% for previous systems (1 of 41)Tavus Griffin postWhether the comparison setup was identical
3.83 VideoFDB generation vs 3.92 human referenceTavus press releaseWhether answers were correct
0.43 s video latency on H100sTavus Griffin postLatency on your hardware or at your volume

Does Tavus say this is risky?

Yes, plainly. The post says the same properties that make these models good interfaces "allow them to deceive", and that is the reason given for limiting Griffin-Lite to select trusted testers while disclosure and safety mechanisms are built (read 2026-10-03). Sume has a related post on Tavus's disclosure settings: Tavus disclosure_type controls vs Sume authored caption cues.

How do I disclose AI in a clip I render with Sume?

Put it in the script. Avatar Video renders the avatar speaking the text you provide, and inline captions burn that same spoken text into the final MP4, so a first line such as "This is an AI-generated presenter" is both heard and read (Generate avatar video). The avatar video docs list no automatic on-screen AI label, so treat disclosure as authored content, not a setting.

Use a multi-scene video_inputs plan if you want the disclosure to open the clip as its own beat, and run an avatar-video preview first so someone reviews the first frames before the full render is paid for (Avatar video previews). A fixed clip has one advantage over a live agent here: the disclosure line is reviewed once and then identical for every viewer.

  • Open with the disclosure line, in speech and in captions.
  • Review the preview stills before the full render.
  • Keep the approved script with the clip so the disclosure is auditable.
  • Do not rely on a vendor's detection claim to protect viewers; assume a clip can be reposted without context.

What would a stronger test look like?

The third-party analysis lists what the evaluations leave open: long conversations, accent coverage and failure rates across varied conditions (kingy.ai, read 2026-10-03). A stronger test would use longer calls, more people recruited outside the vendor, mixed accents and lighting, and a second score for whether the content of the answers was right, not only how human the delivery looked.

For your own product the equivalent is to test with your actual audience. If you plan to ship a rendered presenter, show a handful of people the preview and ask what they think they are watching and whether they would act on what it says. That takes minutes with a preview still and costs less than learning it from comments after launch.

Should a one-minute pass rate change what I build?

It should change how you describe it. If you build a live agent, say it is an agent at the start of the call. If you render a clip, say so in the clip. If your audience will watch only one minute, as in the study, the opening seconds are where that line belongs.

For where Sume sits next to a live agent on latency and model shape, see Tavus Griffin 0.43 s latency vs a Sume async avatar job.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume