Griffin-Lite 26 of 54: error bars and a viewer test for avatars
Tavus says 48% of 54 people took Griffin-Lite for real. With 54 viewers the 95% range is 35% to 61%. A Python check and a test plan for Sume avatar clips.
Tavus reports that 48% of 54 participants thought Griffin-Lite was a real person, against 2.4% of 41 for its earlier Phoenix-4.5 model. With a sample that small the plausible range is wide: a 95% Wilson interval for 26 of 54 is about 35% to 61%, and for 1 of 41 it is about 0.4% to 12.6%. The gap between the two models is large enough to see through the error bars, but the exact 48% should not be quoted as a precise number.
What the page says
The Tavus post, announced October 1, describes a research preview for select trusted testers rather than a customer release. It also lists naturalness of 5.4 out of 7 and trust of 5.6 out of 7. I did not find the survey method beyond the sample sizes, so I do not know how the clips were chosen or shown.
| Model | Rated real | Sample | Share | 95% Wilson interval (my calculation) |
|---|---|---|---|---|
| Griffin-Lite | 26 | 54 | 48.1% | 35.4% to 61.1% |
| Phoenix-4.5 | 1 | 41 | 2.4% | 0.4% to 12.6% |
Compute it yourself
The interval is a few lines of Python. Use it on your own test results.
import math
def wilson(k, n, z=1.96):
p = k / n
d = 1 + z * z / n
c = (p + z * z / (2 * n)) / d
h = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
return round(100 * (c - h), 1), round(100 * (c + h), 1)
print(wilson(26, 54))
print(wilson(1, 41))
print(wilson(20, 40))
Why the sample size matters
The width of an interval shrinks with the square root of the sample. Quadrupling the viewers halves the margin, so going from 54 viewers to about 216 cuts a margin of roughly 13 points to roughly 7. Tavus also compares two different models on two different groups, 54 and 41 people, so the comparison is between groups, not the same viewers rating both. A within-person design, where each viewer sees both versions, would give a tighter answer for the same effort. Whichever you pick, decide the question and the number of viewers before you look at results.
A small viewer test for Sume clips
You can run the same kind of test on avatar clips from Sume's Avatar 1.0. The point is to learn how your viewers react, not to copy the Tavus result.
- Render 5 clips from one script with different avatar handles, plus 5 real recordings of the same lines.
- Show each viewer a mix in random order and ask one question: real person or generated?
- Collect at least 50 viewers; with 50 the margin at a 50% result is about 14 points, and with 100 about 10.
- Report the interval, not only the share.
- Add a question on trust, and keep the answers private to the group.
What Sume does not do
Sume does not publish a realism score for Avatar 1.0, and I found none in the docs. It does not claim that clips pass as real, and viewers should be told when a person is generated. A test like this can help you decide where an avatar is fine and where a real face is better.
Sources
Related posts
More in Sume Avatar 1.0
- Can you buy Tavus Griffin-Lite? Presenter video options today
Tavus Griffin-Lite is a research preview for trusted testers, not a product customers can buy. What Sume ships for presenter videos in the meantime.
- Vidu Q4 audio references vs a scripted avatar: which keeps a voice?
Vidu Q4 Preview takes up to 3 audio references; Sume does not list Vidu. How a scripted avatar video keeps a presenter's face and words consistent.
- What Sume Avatar 1.0 does not do: eight limits to check first
No streaming, no interruption, English-only speech in code, 720p, 4 to 60 seconds, one avatar per video. The limits of Sume Avatar 1.0 in one table.
- YouTube AI disclosure: which avatar pipeline steps are exempt?
YouTube exempts scripts, captions and upscaling from its AI label but not realistic synthetic people. Map each step of a Sume avatar pipeline to the rule.
Written by Sume