Griffin-Lite 26 of 54 Turing result: what the sample size says

Tavus reports 26 of 54 callers fooled by Griffin-Lite after a one-minute call. A 95% interval is about 35% to 61%. Python computes it, and says what to claim.

5 min readSume
All posts

Tavus reports that 48% of participants (26 of 54) judged Griffin-Lite to be a human after a one-minute video call, against 2.4% (1 of 41) for the prior system. With a sample of 54, a 95% Wilson interval for the first result runs from about 35% to 61%, so the honest reading is "roughly a third to three fifths", not "48%". The numbers are from Tavus: Griffin, read 2026-10-05.

That matters if you quote the result. The same page says the Turing test took place after a one-minute video call, and that the model is a research preview for select testers. Quote the counts, not just the percentage, and say what was tested.

Counts and intervals

The table gives the two reported counts and the interval for each, computed with the Wilson score method at 95% confidence. The intervals are my calculation from the counts on the page, not figures that Tavus published.

Reported Turing test results and Wilson 95% intervals, read 2026-10-05
SystemReportedFooled / totalInterval, computed
Griffin-Lite48%26 / 5435.4% to 61.1%
Prior system2.4%1 / 410.4% to 12.6%

Compute it yourself

The code below reproduces both rows. It uses only the standard library, so it runs as written.

from math import sqrt

def wilson(k, n, z=1.96):
    p = k / n
    d = 1 + z * z / n
    c = (p + z * z / (2 * n)) / d
    h = z * sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return c - h, c + h

for name, k, n in [("Griffin-Lite", 26, 54), ("Prior system", 1, 41)]:
    lo, hi = wilson(k, n)
    print("%s: %d/%d = %.1f%%, 95%% CI %.1f%% to %.1f%%" % (name, k, n, 100 * k / n, 100 * lo, 100 * hi))

What you can and cannot claim

Two conclusions survive the uncertainty. The two intervals do not overlap, since the lower end for Griffin-Lite is about 35% and the upper end for the prior system is about 13%, so the gap is real at this sample size. And the interval for Griffin-Lite is wide, which means a figure in the 30s or 60s is as consistent with the data as 48%.

A third point is not about statistics. The test measures whether people took the avatar for a human after one minute. It does not show that the avatar is helpful, accurate or safe, and it says nothing about calls that last longer.

The other numbers on the page

Tavus also reports a generation track score of 3.83 out of 5 on VideoFDB, 1.03 above the next best system at 2.80, and 0.09 below the human reference of 3.92. It reports a perception track score of 3.73, 0.47 below a human reference of 4.20, and an average audio-to-video latency of 0.43 seconds on H100 GPUs. These are Tavus's own figures on its own page; treat them as vendor-reported, and note that I have not reproduced any of them.

If you want a single sentence for a slide, use something like: in a one-minute call, 26 of 54 participants judged Griffin-Lite to be human, which is 48%, with a 95% interval of roughly 35% to 61%. That sentence keeps the counts, the test length, and the uncertainty. It also leaves out claims that the page does not make, such as that the avatar would pass in a longer call or in a different group of people.

How to cite it

If you use this result in a talk or a pitch, give the counts, the one-minute length, and the fact that it is a vendor-reported research preview. If your own avatar is a rendered clip, which viewers watch rather than talk to, a different question applies: whether viewers know it is synthetic. The disclosure post takes that up.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume