Can an AI avatar see me on a video call? Griffin perception vs Sume
Tavus says Griffin perceives visual context: 3.73 of 5 on VideoFDB perception vs 4.20 for humans. Sume avatar clips cannot see you, but a clip can be inspected.
A live avatar can see you only if the model takes in your camera. Tavus says Griffin perceives visual context, and reports 3.73 out of 5 on NVIDIA's VideoFDB perception track against 4.20 for human references. A Sume avatar clip cannot see anyone: it is a render from a script. What Sume can do is inspect a recorded clip after the fact.
What Tavus reports about seeing
The Griffin page lists perception of visual context beyond audio, gaze and gesture control, emotion modelling and an understanding of pauses and silences. On the VideoFDB perception track it reports 3.73 out of 5, which it says is 0.29 ahead of the strongest baseline and 0.47 below the human baseline of 4.20.
Read that gap honestly. By Tavus's own table the model is better than baselines at noticing, and still short of a person. Griffin-Lite is also available only to select trusted testers.
What a rendered avatar knows
Sume's POST /v1/avatar-1.0/talking-video takes a script, an avatar handle and optional product and scene inputs. Nothing from a viewer's camera goes in. Personalisation happens before the render, by writing a different script for each recipient.
Sume does have a way to look at a clip that already exists. POST /v1/video-inspect reads one media.sume.com clip your workspace owns and returns probe facts, sampled stills and an optional transcript. It runs sync by default and waits up to 30 seconds, then returns a queued job if it is not done.
| Question | Griffin-Lite (Tavus) | Sume |
|---|---|---|
| Watches a live camera | Yes, perceives visual context | No |
| VideoFDB perception score | 3.73 of 5 (human 4.20) | Not applicable |
| Reads a recorded clip | Not described on the page | video-inspect: probe, stills, optional transcript |
| Personalisation | In the moment | Per-recipient script before render |
| Availability | Select trusted testers | Public API |
A workable loop
If you want reactive behaviour without a live model, build a loop outside the render. Record a viewer's reply, transcribe or inspect it, write the next script, render the next clip. It is slower than a call, but each step is a job with a result you can store.
See what Griffin changes and what Sume does for the longer comparison.
Sources
Related posts
More in Comparisons
- Can gpt-live-transcribe take an uploaded file? Endpoints vs Sume STT
OpenAI lists gpt-live-transcribe at $0.017/min and only for the realtime transcription endpoint. For a file, Sume STT is $0.01/min, up to 10 minutes per job.
- Cartesia Ink at $0.39 an hour vs ElevenLabs Scribe v2 at $0.22
Cartesia lists Ink STT at $0.39 an hour on the Scale plan; ElevenLabs lists Scribe v2 at $0.22. Ink costs 1.77x as much; Sume's STT is $0.01 a minute.
- Clean vs verbatim transcripts: MAI-Transcribe-2 style vs Sume captions
MAI-Transcribe-2 batch has transcribeStyle clean or verbatim. Sume STT has no style flag; for polished captions supply script_text and keep your wording.
- Clef or Clef-flash in front of a Sume agent: which tier to use
Cloudflare lists a median 38.8 ms for Clef-flash and 209.3 ms for Clef. Use the fast tier to gate runs and the larger one for the choices that cost real money.
Written by Sume