AI video Turing test: how to read Griffin's 48% study
Tavus says 48% of 54 people took Griffin for a real person after a one-minute call. What the study shows, and how to disclose AI in clips you render.

Tavus reports that 48% of people, 26 of 54, believed Griffin was a real person after a one-minute video call, against 2.4% (1 of 41) for previous systems (Tavus Griffin post, read 2026-10-03). It is a vendor-run study with a short call and a small group, so read it as a signal about conversational behavior, not as proof that a video call can no longer be trusted.
The more useful lesson for anyone shipping avatar video, live or rendered, is that audiences may not be able to tell, so disclosure has to come from you.
What did the Tavus study actually measure?
The numbers in the Tavus post are narrow. The Turing-style result is one-minute calls. The separate benchmark result is a 3.83 generation score on NVIDIA's VideoFDB, against a human reference of 3.92 (press release, read 2026-10-03). The Tavus post also reports a perception score of 3.73 and 0.43 seconds of video latency on H100s.
A third-party analysis that we cite only as reported notes the protocol is vendor-conducted and says the evaluations cover conversational realism, not accurate advice, reliable account actions or prompt-injection resistance (kingy.ai, read 2026-10-03). That is the right framing: passing for human for a minute says nothing about whether the agent gave a correct answer.
| Reported figure | Source | What it does not tell you |
|---|---|---|
| 48% believed it was a real person (26 of 54) | Tavus Griffin post | Anything about calls longer than one minute |
| 2.4% for previous systems (1 of 41) | Tavus Griffin post | Whether the comparison setup was identical |
| 3.83 VideoFDB generation vs 3.92 human reference | Tavus press release | Whether answers were correct |
| 0.43 s video latency on H100s | Tavus Griffin post | Latency on your hardware or at your volume |
Does Tavus say this is risky?
Yes, plainly. The post says the same properties that make these models good interfaces "allow them to deceive", and that is the reason given for limiting Griffin-Lite to select trusted testers while disclosure and safety mechanisms are built (read 2026-10-03). Sume has a related post on Tavus's disclosure settings: Tavus disclosure_type controls vs Sume authored caption cues.
How do I disclose AI in a clip I render with Sume?
Put it in the script. Avatar Video renders the avatar speaking the text you provide, and inline captions burn that same spoken text into the final MP4, so a first line such as "This is an AI-generated presenter" is both heard and read (Generate avatar video). The avatar video docs list no automatic on-screen AI label, so treat disclosure as authored content, not a setting.
Use a multi-scene video_inputs plan if you want the disclosure to open the clip as its own beat, and run an avatar-video preview first so someone reviews the first frames before the full render is paid for (Avatar video previews). A fixed clip has one advantage over a live agent here: the disclosure line is reviewed once and then identical for every viewer.
- Open with the disclosure line, in speech and in captions.
- Review the preview stills before the full render.
- Keep the approved script with the clip so the disclosure is auditable.
- Do not rely on a vendor's detection claim to protect viewers; assume a clip can be reposted without context.
What would a stronger test look like?
The third-party analysis lists what the evaluations leave open: long conversations, accent coverage and failure rates across varied conditions (kingy.ai, read 2026-10-03). A stronger test would use longer calls, more people recruited outside the vendor, mixed accents and lighting, and a second score for whether the content of the answers was right, not only how human the delivery looked.
For your own product the equivalent is to test with your actual audience. If you plan to ship a rendered presenter, show a handful of people the preview and ask what they think they are watching and whether they would act on what it says. That takes minutes with a preview still and costs less than learning it from comments after launch.
Should a one-minute pass rate change what I build?
It should change how you describe it. If you build a live agent, say it is an agent at the start of the call. If you render a clip, say so in the clip. If your audience will watch only one minute, as in the study, the opening seconds are where that line belongs.
For where Sume sits next to a live agent on latency and model shape, see Tavus Griffin 0.43 s latency vs a Sume async avatar job.
Sources
Related posts
More in Comparisons
- Claude Code mod vs MCP server vs skill vs hook: where Sume fits
Claude Code mods, MCP servers, skills and settings hooks overlap. Which one gives an agent Sume's image and video tools, and which one only guards them.
- Comfy Agent Ask or Auto mode vs unattended Sume Format runs
Comfy Agent asks before each run or runs on its own. Sume API runs never ask. See which controls replace the approval prompt when a batch goes overnight.
- Comfy Agent vs an MCP agent for image and video work
Comfy Agent builds and runs ComfyUI graphs for you in Comfy Cloud; an MCP agent calls hosted tools like Sume's. What differs in control, billing and location.
- Compare a local open-weights image model with hosted ones fairly
Same prompt, same shape, no seed: how to test a local Ideogram 4 run against Sume's hosted image models, with a script that prints one image URL per model.
Written by Sume