Can an AI avatar see me on a video call? Griffin perception vs Sume

Tavus says Griffin perceives visual context: 3.73 of 5 on VideoFDB perception vs 4.20 for humans. Sume avatar clips cannot see you, but a clip can be inspected.

5 min readSume
All posts

A live avatar can see you only if the model takes in your camera. Tavus says Griffin perceives visual context, and reports 3.73 out of 5 on NVIDIA's VideoFDB perception track against 4.20 for human references. A Sume avatar clip cannot see anyone: it is a render from a script. What Sume can do is inspect a recorded clip after the fact.

What Tavus reports about seeing

The Griffin page lists perception of visual context beyond audio, gaze and gesture control, emotion modelling and an understanding of pauses and silences. On the VideoFDB perception track it reports 3.73 out of 5, which it says is 0.29 ahead of the strongest baseline and 0.47 below the human baseline of 4.20.

Read that gap honestly. By Tavus's own table the model is better than baselines at noticing, and still short of a person. Griffin-Lite is also available only to select trusted testers.

What a rendered avatar knows

Sume's POST /v1/avatar-1.0/talking-video takes a script, an avatar handle and optional product and scene inputs. Nothing from a viewer's camera goes in. Personalisation happens before the render, by writing a different script for each recipient.

Sume does have a way to look at a clip that already exists. POST /v1/video-inspect reads one media.sume.com clip your workspace owns and returns probe facts, sampled stills and an optional transcript. It runs sync by default and waits up to 30 seconds, then returns a queued job if it is not done.

Seeing the viewer vs inspecting a clip (read 2026-10-05)
QuestionGriffin-Lite (Tavus)Sume
Watches a live cameraYes, perceives visual contextNo
VideoFDB perception score3.73 of 5 (human 4.20)Not applicable
Reads a recorded clipNot described on the pagevideo-inspect: probe, stills, optional transcript
PersonalisationIn the momentPer-recipient script before render
AvailabilitySelect trusted testersPublic API

A workable loop

If you want reactive behaviour without a live model, build a loop outside the render. Record a viewer's reply, transcribe or inspect it, write the next script, render the next clip. It is slower than a call, but each step is a job with a result you can store.

See what Griffin changes and what Sume does for the longer comparison.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume