Griffin-Lite 26 of 54 Turing result: what the sample size says
Tavus reports 26 of 54 callers fooled by Griffin-Lite after a one-minute call. A 95% interval is about 35% to 61%. Python computes it, and says what to claim.

Tavus reports that 48% of participants (26 of 54) judged Griffin-Lite to be a human after a one-minute video call, against 2.4% (1 of 41) for the prior system. With a sample of 54, a 95% Wilson interval for the first result runs from about 35% to 61%, so the honest reading is "roughly a third to three fifths", not "48%". The numbers are from Tavus: Griffin, read 2026-10-05.
That matters if you quote the result. The same page says the Turing test took place after a one-minute video call, and that the model is a research preview for select testers. Quote the counts, not just the percentage, and say what was tested.
Counts and intervals
The table gives the two reported counts and the interval for each, computed with the Wilson score method at 95% confidence. The intervals are my calculation from the counts on the page, not figures that Tavus published.
| System | Reported | Fooled / total | Interval, computed |
|---|---|---|---|
| Griffin-Lite | 48% | 26 / 54 | 35.4% to 61.1% |
| Prior system | 2.4% | 1 / 41 | 0.4% to 12.6% |
Compute it yourself
The code below reproduces both rows. It uses only the standard library, so it runs as written.
from math import sqrt
def wilson(k, n, z=1.96):
p = k / n
d = 1 + z * z / n
c = (p + z * z / (2 * n)) / d
h = z * sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
return c - h, c + h
for name, k, n in [("Griffin-Lite", 26, 54), ("Prior system", 1, 41)]:
lo, hi = wilson(k, n)
print("%s: %d/%d = %.1f%%, 95%% CI %.1f%% to %.1f%%" % (name, k, n, 100 * k / n, 100 * lo, 100 * hi))What you can and cannot claim
Two conclusions survive the uncertainty. The two intervals do not overlap, since the lower end for Griffin-Lite is about 35% and the upper end for the prior system is about 13%, so the gap is real at this sample size. And the interval for Griffin-Lite is wide, which means a figure in the 30s or 60s is as consistent with the data as 48%.
A third point is not about statistics. The test measures whether people took the avatar for a human after one minute. It does not show that the avatar is helpful, accurate or safe, and it says nothing about calls that last longer.
The other numbers on the page
Tavus also reports a generation track score of 3.83 out of 5 on VideoFDB, 1.03 above the next best system at 2.80, and 0.09 below the human reference of 3.92. It reports a perception track score of 3.73, 0.47 below a human reference of 4.20, and an average audio-to-video latency of 0.43 seconds on H100 GPUs. These are Tavus's own figures on its own page; treat them as vendor-reported, and note that I have not reproduced any of them.
If you want a single sentence for a slide, use something like: in a one-minute call, 26 of 54 participants judged Griffin-Lite to be human, which is 48%, with a 95% interval of roughly 35% to 61%. That sentence keeps the counts, the test length, and the uncertainty. It also leaves out claims that the page does not make, such as that the avatar would pass in a longer call or in a different group of people.
How to cite it
If you use this result in a talk or a pitch, give the counts, the one-minute length, and the fact that it is a vendor-reported research preview. If your own avatar is a rendered clip, which viewers watch rather than talk to, a different question applies: whether viewers know it is synthetic. The disclosure post takes that up.
Sources
Related posts
More in Sume Avatar 1.0
- Tavus Video to Face replica vs a Sume avatar from a photo or prompt
Tavus builds a replica from video or a photo; Sume builds an avatar from a prompt, props or a public photo URL. What each input gives you, and what it costs.
- What an AI avatar may not claim on TikTok Shop
TikTok Shop prohibits AI that impersonates real people or invents doctors and experts to endorse products. An avatar can present, but it can't be a fake expert.
- UGC-style ad batch: ten 12-second hooks on Avatar 1.0 for $29.40
Ten 12-second UGC-style hook variants cost about $29.40 on Sume Avatar 1.0 plus, $22.08 on standard. The request, captions and checks.
- UGC-style avatar ad: hook, silent demo beat and CTA in video_inputs
Build a three-scene UGC avatar ad with POST /v1/avatar-1.0/talking-video: a spoken hook, a silent demo beat, a spoken CTA, inside the 4 to 60 second window.
Written by Sume