NVIDIA VideoFDB: what Griffin Lite's 3.73 means for avatar buyers
VideoFDB scores perception on 237 call clips across 11 nonverbal dynamics. Griffin Lite led full-duplex systems at 3.73 against a human 4.20.
NVIDIA's VideoFDB benchmark tests whether an agent perceives what happens in a video call, using 237 dyadic clips from real calls across 11 nonverbal conversational dynamics. Tavus Griffin Lite scored 3.73 overall, the highest overall perception score the project page reports, against 3.40 for the best open-source model, MiniCPM-o 4.5, and 4.20 for the human reference.
For a buyer, the useful reading is narrow: it measures perception in a live conversation, not the quality of a rendered video.
What the benchmark reports
NVIDIA's project page says no evaluated agent approaches the human reference. The largest gap it highlights is conversational flow: humans score 4.20, closed-source models 2.20-2.81, and the best open-source model 3.54. It also reports that audio-visual input scored worse than audio-only on perception rubrics in every model family, naming captioning collapse and visual-stream ignorance as dominant failure modes (NVIDIA Research: VideoFDB).
| Item | Value |
|---|---|
| Clips | 237 dyadic clips from real video calls |
| Nonverbal dynamics | 11 |
| Griffin Lite overall perception | 3.73 |
| Best open-source model (MiniCPM-o 4.5) | 3.40 |
| Human reference | 4.20 |
What it does not tell a buyer
VideoFDB scores how well an agent reads a human on a call. It does not score lip sync, voice quality, price or latency, and it does not cover scripted video. Tavus describes Griffin as a full-duplex model; its own page is the place to check availability before planning around it (Tavus Griffin).
If your use is a message sent to many people, perception is not the property you need. You need predictable words, a stable face, captions and a known cost per clip. Sume's avatar video takes a script and renders a file, with a 4-60 second window per job (Generate avatar video).
- Use the benchmark to question vendors about live perception, not to rank rendered clips.
- Ask for the specific dynamics your use depends on, such as interruptions or backchannels.
- For a rendered clip, judge a sample file at your target aspect ratio and length.
Questions for a live-agent vendor
Ask which of the 11 dynamics the agent handles worst and how that shows up in your use case. Ask whether the vendor measured with visual input on and off, given that VideoFDB reports visual input lowered perception scores.
Ask for a live trial in your own conditions: lighting, accents, background noise. A benchmark score is an average over its clips, not a promise about your callers.
Reading the numbers with care
A score of 3.73 against 3.40 is a difference between two systems on one benchmark. It does not say how either behaves on your calls. The gap to the human reference of 4.20 is the larger fact: the benchmark's own conclusion is that no evaluated agent approaches human perception.
Treat the result as a reason to keep a human path for anything high stakes, and as a prompt to ask vendors what they have measured themselves.
Sources
Related posts
More in Models
- OmniVoice: 600+ languages, CC-BY-NC weights, hosted TTS instead
OmniVoice covers 600+ languages in a 0.6B model, but its weights are CC-BY-NC. What the card says, what it omits, and where a hosted TTS job fits.
- Pick an ElevenLabs model by language count: 90+, 70+, 32, 29
ElevenLabs lists 90+ languages for v4, 70+ for v3, 32 for Flash v2.5, 29 for Multilingual v2 and English only for Flash v2. Where Sume Sonic fits.
- Pocket TTS speed: 6x on an M4 CPU, 2.3-2.5x on a VM CPU
Kyutai Pocket TTS runs about 6x real time on an M4 CPU and 2.3-2.5x on a 4-vCPU cloud VM. It cannot add silence for pauses. What to plan for.
- Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.
Written by Sume