Griffin-Lite 26 of 54: error bars and a viewer test for avatars

Tavus says 48% of 54 people took Griffin-Lite for real. With 54 viewers the 95% range is 35% to 61%. A Python check and a test plan for Sume avatar clips.

5 min readSume
All posts

Tavus reports that 48% of 54 participants thought Griffin-Lite was a real person, against 2.4% of 41 for its earlier Phoenix-4.5 model. With a sample that small the plausible range is wide: a 95% Wilson interval for 26 of 54 is about 35% to 61%, and for 1 of 41 it is about 0.4% to 12.6%. The gap between the two models is large enough to see through the error bars, but the exact 48% should not be quoted as a precise number.

What the page says

The Tavus post, announced October 1, describes a research preview for select trusted testers rather than a customer release. It also lists naturalness of 5.4 out of 7 and trust of 5.6 out of 7. I did not find the survey method beyond the sample sizes, so I do not know how the clips were chosen or shown.

Griffin-Lite and Phoenix-4.5 as reported, with my intervals (read 2026-10-08)
ModelRated realSampleShare95% Wilson interval (my calculation)
Griffin-Lite265448.1%35.4% to 61.1%
Phoenix-4.51412.4%0.4% to 12.6%

Compute it yourself

The interval is a few lines of Python. Use it on your own test results.

import math

def wilson(k, n, z=1.96):
    p = k / n
    d = 1 + z * z / n
    c = (p + z * z / (2 * n)) / d
    h = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return round(100 * (c - h), 1), round(100 * (c + h), 1)

print(wilson(26, 54))
print(wilson(1, 41))
print(wilson(20, 40))

Why the sample size matters

The width of an interval shrinks with the square root of the sample. Quadrupling the viewers halves the margin, so going from 54 viewers to about 216 cuts a margin of roughly 13 points to roughly 7. Tavus also compares two different models on two different groups, 54 and 41 people, so the comparison is between groups, not the same viewers rating both. A within-person design, where each viewer sees both versions, would give a tighter answer for the same effort. Whichever you pick, decide the question and the number of viewers before you look at results.

A small viewer test for Sume clips

You can run the same kind of test on avatar clips from Sume's Avatar 1.0. The point is to learn how your viewers react, not to copy the Tavus result.

  • Render 5 clips from one script with different avatar handles, plus 5 real recordings of the same lines.
  • Show each viewer a mix in random order and ask one question: real person or generated?
  • Collect at least 50 viewers; with 50 the margin at a 50% result is about 14 points, and with 100 about 10.
  • Report the interval, not only the share.
  • Add a question on trust, and keep the answers private to the group.

What Sume does not do

Sume does not publish a realism score for Avatar 1.0, and I found none in the docs. It does not claim that clips pass as real, and viewers should be told when a person is generated. A test like this can help you decide where an avatar is fine and where a real face is better.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume