What is a full-duplex AI avatar, and when is a rendered clip enough?

A full-duplex avatar listens and talks at once, as in Tavus Griffin-Lite. If your viewers never talk back, a rendered Sume avatar clip does the job.

5 min readSume
All posts

A full-duplex AI avatar listens and speaks at the same time, the way two people on a video call do, instead of waiting for you to finish before it replies. If your viewers only watch and never talk back, you do not need one: a rendered avatar clip says the same approved script every time, and Sume Avatar 1.0 makes those.

Half-duplex versus full-duplex

Most avatar products so far are turn-based. A voice model detects that you stopped talking, an LLM writes a reply, a speech model says it, and an avatar renderer animates it. Each hand-off adds delay and the avatar cannot react while you speak.

Tavus describes Griffin as a model that never stops thinking: every sub-second mini-turn it decides whether to speak up, react or wait. It takes in audio and video continuously and produces speech and video together. That is what full-duplex means in practice: nodding while you talk, being interrupted, and starting a sentence over.

What Tavus publishes about Griffin

The Griffin page lists an average audio-to-video latency of 0.43 seconds on H100 GPUs, 720p video in 320 ms chunks, and audio packets as small as 10 ms. It reports that 48 percent of participants believed Griffin was a real person after a one-minute video call, and a score of 3.83 out of 5 on NVIDIA VideoFDB against a human reference of 3.92.

Griffin-Lite is the research preview and is available to a select group of early testers. The page says it will not be available for customers at this time and gives no pricing.

Live conversation versus a rendered clip (Tavus figures read 2026-10-05)
QuestionFull-duplex avatar (Griffin-Lite preview)Rendered clip (Sume Avatar 1.0)
Can it react while the viewer talks?Yes, by designNo
Can you buy it today?No, select testers onlyYes, through API, MCP and the dashboard
Does every viewer get the same words?No, it improvisesYes, you approve the script first
LengthOpen-ended call4 to 60 seconds per job
Cost grows with viewers?Yes, per live minuteNo, one render, many plays

Signs a rendered clip is enough

  • The message is the same for every viewer: an onboarding step, a product update, a policy notice.
  • Someone has to approve the exact wording before it ships.
  • The video lives in an email, a page or a feed, not in a call.
  • You want to check the first frame and pay only after you like it.

Signs you need live conversation

For the first group, Sume lets you preview the first frame with POST /v1/avatar-video-previews, adjust it, then generate the final video from that preview. A job is one avatar for one video, 4 to 60 seconds, at 720p, with aspect ratios from 1:1 to 16:9.

  • The viewer asks questions you cannot list in advance.
  • The value is practice or role play, where the avatar must respond to what was said.
  • A person expects an answer in about a second.

What to do

Write down the three questions a viewer might ask after watching. If you can answer all three in a follow-up clip, render the clips. If the questions are open-ended, keep the live product for the conversation and use rendered clips for the parts that never change.

Sources

Related posts

More in Models

All Models posts

Written by Sume