Griffin VideoFDB 3.83 vs human 3.92: what the scores cover
Tavus reports Griffin-Lite at 3.83 of 5 on VideoFDB generation against a 3.92 human reference. Those scores rate live conversation, not a scripted ad clip.

Tavus's Griffin page reports that Griffin-Lite scores 3.83 out of 5 on the VideoFDB generation track, against a human reference of 3.92, and 3.73 on the perception track against a human 4.20 (Tavus Griffin page, read 2026-10-06). Those numbers rate a model that listens and answers in real time. They say nothing about how a scripted, rendered clip will look, so they are not a score you can reuse to pick an avatar for an ad.
What the page says about the scores
The page lists two VideoFDB tracks and a human reference for each. Tavus says Griffin-Lite leads the next system by 1.03 points on generation and 0.29 points on perception.
Griffin-Lite is a research preview for select trusted testers. The page says the full Griffin is not yet available and gives no pricing, so there is no API call to try the model with today.
| Track | Griffin-Lite | Human reference |
|---|---|---|
| Generation | 3.83 | 3.92 |
| Perception | 3.73 | 4.20 |
Why a live-conversation score does not transfer to a clip
A full-duplex score judges turn-taking, listening cues and reaction in a live exchange. A script-driven clip has no listener, so none of those behaviours are in play. What you judge on a clip is the face, the mouth and the pacing of lines you wrote.
Sume ships the script-driven side. Generate avatar video takes a ready avatar plus either a script or ordered video_inputs, for an estimated 4 to 60 seconds, and returns a rendered MP4 at 720p.
How to judge a Sume clip yourself
Use the first-frame preview to approve the composition before paying for the render. Avatar video previews create stills only, and generate-video on the preview id starts the full render. The stills are tier-independent, so you can approve and then change quality without a new preview.
Then watch the finished clip with sound on a phone, which is where short ads are seen. A score table cannot do that for you, and Sume does not publish a VideoFDB number for Avatar 1.0.
Bottom line
Treat VideoFDB as a signal about live avatars from a vendor whose model you cannot call yet. If your job is a scripted spokesperson, the choice is made by the checks you can run on a rendered clip: lip movement, framing, pacing and cost.
Sources
Related posts
More in Sume Avatar 1.0
- Gym class schedule change: a 20-second avatar video, priced
A 20-second avatar notice for a changed class time costs about $3.68 on Standard, $4.90 on Plus or $11.00 on Max at Sume's listed per-second rates.
- HOA community notice: a 30-second avatar video for residents, priced
A 30-second avatar video for an HOA or building notice costs about $5.52 on Standard, $7.35 on Plus or $16.50 on Max at Sume's listed per-second rates.
- Interview video with one AI avatar host and guest cards on screen
A final avatar video holds one avatar, so an interview needs a host clip plus guest stills. Use Timeline compose at $0.02 a shot to put a still beside the host.
- Korean talking avatar captions: black-outline or korean-ad, not slam
Korean script on a Sume avatar video with slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style. Use black-outline or korean-ad.
Written by Sume