Interactive avatar or avatar video? A five-question test

Griffin-Lite is a research preview. Five questions tell you whether you need a live avatar or a rendered Sume Avatar 1.0 clip, and what to ship this quarter.

4 min readSume
All posts

Choose an interactive avatar only if the person on screen must react to what the viewer says or does while the viewer is still there. If you can write the words ahead of time, a rendered avatar video is the simpler product to ship now. Sume Avatar 1.0 is the rendered kind: you submit a script, poll a job, and get an MP4 back.

The question is timely because Tavus announced Griffin-Lite on 1 October 2026. It is a research preview for select testers, so most teams cannot build on it yet. The five questions below work for any live-avatar vendor, not just this one.

What was announced, and what is available

Tavus describes Griffin as a Human Interaction Model that understands and generates face-to-face real-time interaction, as one video-to-video system. Griffin-Lite is available to select trusted testers as a research preview, and Tavus says further alignment and safety procedures are required for a safe release, and that Griffin-Lite is not available to customers at this time.

Griffin-Lite claims on the vendor page (read 2026-10-07)
ItemWhat the page says
AvailabilityResearch preview, select trusted testers via a request form
Turing-style test48% (n=54) believed their partner was a real person after a one-minute video call
Average latency0.43 seconds for video generation
Suggested usesTutoring, professional conversation practice, visual problem-solving

The five questions

Answer them in order. One strong yes to question 1 or 2 points to a live avatar. Otherwise a rendered clip is enough.

  • 1. Does the avatar need to answer a question nobody could predict? Open-ended tutoring and support calls do. A product explainer does not.
  • 2. Must it perceive the viewer, such as their face or tone? A rendered clip cannot see anyone.
  • 3. Is the same video shown to many people? If yes, render it once and reuse it. Live generation per viewer is wasted work.
  • 4. Does someone have to approve the output before it reaches a customer? Review needs a finished file, which a live session never produces.
  • 5. Do you need it this quarter? A research preview has no public date, so plan around what has an API today.

What the rendered path gives you

Avatar 1.0 starts with a reusable avatar made from a prompt, structured traits, or a reference photo. You then call POST /v1/avatar-1.0/talking-video with that avatar handle and a script. Sume accepts scripts whose estimated length is 4 to 60 seconds, and returns a job you poll until it is completed (Generate avatar video).

  • Aspect ratios 1:1, 3:4, 9:16, 4:3 and 16:9; resolution is 720p at this time.
  • Quality standard, plus (default) or max.
  • Optional inline captions burned into the final MP4.
  • A preview stage so you can approve the first frame before a full render.

Where Sume does not fit

Sume does not document a live, two-way avatar. There is no streaming session, and a job is accepted, queued, processed and fetched (Jobs and results). If your answer to question 1 or 2 was a firm yes, a rendered clip will not replace a live conversation.

A common middle path is to write a small set of scripted clips for the questions you do expect, and pick between them with your own logic. That is closer to a help-center video library than a conversation, and it can ship without waiting for any preview.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume