Full-duplex video AI: what Griffin changes, what Sume does
Full-duplex video AI listens, watches and answers at once. Tavus Griffin is gated; Sume makes scripted avatar clips by job. Where each fits.

Full-duplex video AI is a model that listens, watches and speaks at the same time, instead of waiting for you to finish a sentence and then answering. Tavus calls its version Griffin and describes it as a video-to-video model that "sees, hears, speaks and moves at the same time, like a person" (read 2026-10-03). Sume does not have this: Sume renders avatar video as a job, from a script you write, and hands back a finished clip rather than a live stream.
This post explains what full-duplex changes technically, what you can actually use today, and which kind of project each approach suits.
What does full-duplex mean for a video agent?
A turn-based avatar waits for silence, then answers. Tavus describes Griffin differently: it "never stops thinking", and every sub-second mini-turn it decides whether to speak up, react or wait while it keeps listening and watching. In the same description it can interrupt, backchannel, nod and shift emotional tone mid-conversation (Tavus Griffin post, read 2026-10-03).
The practical result is that the small things people do in a call, such as a nod, an "mm-hm", or stopping when you cut in, come from the model itself instead of from rules bolted on top. That is what the word full-duplex is pointing at, borrowed from telephony where both parties can talk at once.
How is that different from a cascaded avatar pipeline?
Most live avatars today are a chain. Anam's documentation describes a persona as a face, a voice, an LLM and a system prompt, and lists four stages: speech-to-text captures the user, an LLM writes the reply, text-to-speech makes audio, and face generation draws the speaker (Anam docs, read 2026-10-03). Tavus's own platform page lists separate models for rendering (Phoenix-4.5, 134 ms latency), perception (Raven-1, under 300 ms context freshness) and turn-taking (Sparrow-2), as the same page states.
Griffin, per Tavus, replaces that hand-off chain with one model. Tavus reports 0.43 seconds of video latency on H100s and 3.83 on NVIDIA's VideoFDB generation score, against a human reference of 3.92 (Tavus Griffin post and press release, read 2026-10-03). Those are vendor-reported figures on a vendor-chosen setup.
| Property | Cascaded live avatar | Tavus Griffin | Sume Avatar Video |
|---|---|---|---|
| Interaction | Turn by turn, four stages (Anam docs) | Continuous, sub-second mini-turns (Tavus) | None: one script in, one clip out |
| Output | A live video stream | A live video stream | An MP4 of 4-60 seconds |
| Who can use it | Offered by vendors such as Anam and Tavus | Griffin-Lite: select trusted testers only (Tavus) | Any workspace with an API key |
| Reacts to the viewer | Yes, via the pipeline | Yes, natively (Tavus) | No |
| Review the output before viewers see it | Not possible in a live stream | Not possible in a live stream | Yes, preview stills first |
Can I use Griffin today?
Not as a customer. Tavus writes that Griffin-Lite "will not be available for use for customers at this time, though it is available for select trusted testers as a research preview", and ties broader release to extra disclosure and safety mechanisms (Tavus Griffin post, read 2026-10-03). A third-party analysis, which we cite only as reported, found no public Griffin price or release date (kingy.ai, read 2026-10-03).
So the useful question for a team this week is not "when do we switch" but "which of our conversations need to be live at all". Tavus's own examples are tutoring, rehearsing difficult conversations and product troubleshooting (press release). Those are interactive by nature. A product announcement, a welcome message or a how-to is not.
What does Sume do instead?
Sume has no live session, no WebRTC room and no listening step. Avatar Video 1.0 takes a ready avatar handle plus either a script or ordered video_inputs, estimates the duration, accepts 4-60 seconds, and returns a job you poll or receive by webhook (Generate avatar video).
What it gives you that a live agent does not is control. You can run avatar-video-previews to approve first frames before paying for the full render, pick standard, plus or max quality, add silence beats between spoken scenes, and burn captions in. The same clip then plays identically for every viewer, which is the point if the message is a fixed one.
- Choose a live full-duplex agent when the viewer's next words change what the avatar should say.
- Choose a rendered clip when you can write the answer in advance and want to review it once.
- Use both: a live agent for the open conversation, and rendered clips for the answers you give every time.
What does it mean for latency and cost planning?
Latency is the headline number for live agents, and Tavus leads with it: 0.43 seconds of video latency on H100s, which it describes as half that of the next fastest method (Tavus Griffin post, read 2026-10-03). The same figure also hints at the hardware class involved. Streaming video from a model that runs continuously needs a GPU dedicated to each conversation, which is why live plans are metered by conversation minute and limited by concurrent streams.
A rendered clip works the other way round. Sume does not stream frames, so there is no per-viewer GPU. The cost is paid when the job runs, and the viewer later plays an ordinary video file. That is why the same sentence delivered to ten thousand people is a single render, and why a conversation that must differ for each of them is not a good fit for Sume.
How do I check which side my project is on?
Write down the first three things a viewer might say. If the avatar's reply to each is the same sentence no matter who is asking, you have a clip. If the reply depends on what they just did, showed on camera or interrupted with, you have a conversation, and a clip will feel scripted because it is.
For a head-to-head on session API shape, see Real-time AI avatar vs video avatar API. For what to render while Griffin is gated, see Tavus Griffin-Lite is not available: what to render today.
Sources
- Tavus: Introducing Griffin (read 2026-10-03)
- Tavus press release, via Yahoo Finance (read 2026-10-03)
- Tavus home page (read 2026-10-03)
- Anam docs: introduction (read 2026-10-03)
- kingy.ai: Tavus Griffin specs, benchmarks, availability (third-party analysis, read 2026-10-03)
- Generate avatar video
- Jobs and results
Related posts
More in Comparisons
- GEMA v Suno ruling: what to check before AI music goes in an ad
A Munich court ruled against Suno on 31 July 2026, not final. What it says, what it leaves open, and a record to keep for any AI track you put in an ad.
- Pocket TTS voice cloning: a wav in, and what Sume does instead
Pocket TTS clones from a wav file you pass to --voice, with consent rules in its model card. Sume's API takes voice ids, not audio. Here is the difference.
- Schedule a weekly AI video: Sume Scheduled, Hermes cron or API
Three ways to run an AI video on a weekly clock with Sume: a dashboard schedule, a Hermes cron job, or plain cron calling a Format. What each can and cannot do.
- Which image and video models on Sume have open weights?
Sume's catalog is hosted. Where vendors publish weights for families Sume lists (Qwen-Image, MiniMax H3, FLUX.2) and why a null hugging_face_id proves nothing.
Written by Sume