NVIDIA VideoFDB: what Griffin Lite's 3.73 means for avatar buyers

VideoFDB scores perception on 237 call clips across 11 nonverbal dynamics. Griffin Lite led full-duplex systems at 3.73 against a human 4.20.

5 min readSume
All posts

NVIDIA's VideoFDB benchmark tests whether an agent perceives what happens in a video call, using 237 dyadic clips from real calls across 11 nonverbal conversational dynamics. Tavus Griffin Lite scored 3.73 overall, the highest overall perception score the project page reports, against 3.40 for the best open-source model, MiniCPM-o 4.5, and 4.20 for the human reference.

For a buyer, the useful reading is narrow: it measures perception in a live conversation, not the quality of a rendered video.

What the benchmark reports

NVIDIA's project page says no evaluated agent approaches the human reference. The largest gap it highlights is conversational flow: humans score 4.20, closed-source models 2.20-2.81, and the best open-source model 3.54. It also reports that audio-visual input scored worse than audio-only on perception rubrics in every model family, naming captioning collapse and visual-stream ignorance as dominant failure modes (NVIDIA Research: VideoFDB).

VideoFDB headline numbers (read 2026-10-03)
ItemValue
Clips237 dyadic clips from real video calls
Nonverbal dynamics11
Griffin Lite overall perception3.73
Best open-source model (MiniCPM-o 4.5)3.40
Human reference4.20

What it does not tell a buyer

VideoFDB scores how well an agent reads a human on a call. It does not score lip sync, voice quality, price or latency, and it does not cover scripted video. Tavus describes Griffin as a full-duplex model; its own page is the place to check availability before planning around it (Tavus Griffin).

If your use is a message sent to many people, perception is not the property you need. You need predictable words, a stable face, captions and a known cost per clip. Sume's avatar video takes a script and renders a file, with a 4-60 second window per job (Generate avatar video).

  • Use the benchmark to question vendors about live perception, not to rank rendered clips.
  • Ask for the specific dynamics your use depends on, such as interruptions or backchannels.
  • For a rendered clip, judge a sample file at your target aspect ratio and length.

Questions for a live-agent vendor

Ask which of the 11 dynamics the agent handles worst and how that shows up in your use case. Ask whether the vendor measured with visual input on and off, given that VideoFDB reports visual input lowered perception scores.

Ask for a live trial in your own conditions: lighting, accents, background noise. A benchmark score is an average over its clips, not a promise about your callers.

Reading the numbers with care

A score of 3.73 against 3.40 is a difference between two systems on one benchmark. It does not say how either behaves on your calls. The gap to the human reference of 4.20 is the larger fact: the benchmark's own conclusion is that no evaluated agent approaches human perception.

Treat the result as a reason to keep a human path for anything high stakes, and as a prompt to ask vendors what they have measured themselves.

Sources

Related posts

More in Models

All Models posts

Written by Sume