Griffin-Lite's 0.43 s latency: what a recorded avatar clip gives up

Tavus reports Griffin-Lite video latency of 0.43 s on average (0.27-0.59 s) in a research preview. A recorded clip on Sume is a job, so choose by use case.

5 min readSume
All posts

Tavus says Griffin-Lite generates video with an average latency of 0.43 seconds on H100 GPUs, with a best case of 0.27 seconds and a worst case of 0.59 seconds, in the research preview it announced on 1 October 2026 (Tavus Griffin page, read 2026-10-09). That is the speed of a live, two-way video conversation. A recorded avatar clip from Sume is a different product: you send a script, you get a finished MP4 from a job, and nobody is waiting on the other end of a call.

The question that decides between them is simple: does a person need the answer while they are still looking at the screen?

What the vendor page says

The Tavus page describes Griffin-Lite as a full-duplex, video-to-video model: it perceives a person's audio and video and generates its own face and voice together. It says the next fastest competitor it tested, AvatarForcing, averaged 0.86 seconds. It also says Griffin-Lite is available only to a select group of early testers and will not be available to customers at this time, and that Tavus is working on safe disclosure features first. The page lists no price.

Griffin-Lite facts from the Tavus page (read 2026-10-09)
ItemWhat the page says
StatusResearch preview, select early testers
Customer availabilityNot available to customers at this time
Average video latency (H100)0.43 s
Best / worst case0.27 s / 0.59 s
Competitor average (AvatarForcing)0.86 s
PriceNot listed

What Sume offers instead

Sume's docs describe avatar video as a job. You POST a script to /v1/avatar-1.0/talking-video, receive a job id, and poll GET /v1/jobs/:id/status until it completes, then fetch /result. The docs list no live video-call endpoint, so for a real-time conversation Sume is not the tool today. The sync mode on submits waits at most 30 seconds before returning, which is a bounded wait, not a stream.

What you get in exchange is a durable file you can review, caption and reuse. Price is $0.184 per second at standard, $0.245 at plus and $0.55 at max, so a 20-second greeting is $3.68, $4.90 or $11.00 (catalog, read 2026-10-09).

A decision rule

Use a live model when the viewer talks back and the reply must adapt: a support call, a tutoring session, a mock interview. Use a recorded clip when the message is known in advance and you want to approve it: an onboarding welcome, a product update, a localized ad.

Many teams need both. They script the welcome as a recorded clip today and keep a waitlist entry for a live agent. That plan avoids waiting on a preview that, per the vendor page, has no customer date. See how to read the 48 percent study before you quote it in a deck.

  • Live conversation: waitlist for Griffin-Lite, no price published.
  • Known script: Sume avatar video, billed per output second.
  • Both: ship the recorded version first.

Latency is not the only number

A latency figure describes how fast the next frame appears. It does not describe cost, access or consent. On the Tavus page, cost is absent and access is limited to early testers. For a recorded clip, the number that matters is not milliseconds but turnaround and spend: you submit, the job runs under your workspace's generation slots, and you fetch the file when it completes. The docs give no fixed turnaround, and standard is described as the fastest Sume execution path while max is slower.

If a launch date depends on a live agent, treat 0.43 seconds as a vendor claim about a preview and plan around the fact that customers cannot use it yet.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume