Tavus Human Interaction Model: what Griffin is, and what ships today
Tavus calls Griffin a Human Interaction Model. Only select testers have Griffin-Lite. Here is what it is, and what ships today on Sume for a talking presenter.

A Human Interaction Model (HIM) is Tavus's name for a model that understands and generates face-to-face, real-time interaction, and Griffin is the first. You cannot buy it today: Griffin-Lite is a research preview for select testers, so a talking presenter you can ship this week has to come from a rendered avatar.
What Tavus says it is
The Griffin page calls it the world's first Human Interaction Model, a new class of model designed to understand and generate face-to-face real-time human interaction. It perceives audio and video continuously and generates speech and video at once. The page lists a custom codec called Tavec, with 40 values per frame at 100 frames per second.
Tavus attributes these results to it: 0.43 seconds average video latency on H100s, 720p video in 320 ms chunks, and first place on the DOVER, FID and THEval visual benchmarks. Those are vendor claims on the vendor's own page.
What is and is not available
Griffin-Lite is available today to a select group of early testers. The page says it will not be available for customers at this time, and it ties a wider release to safety concerns being addressed. No price is published, so any cost comparison today would be invented.
| Item | What the page says |
|---|---|
| Access | Select group of early testers |
| Customer availability | Not at this time |
| Pricing | Not published |
| Output | 720p video in 320 ms chunks |
| Latency | 0.43 seconds average on H100s |
What ships on Sume
Sume Avatar 1.0 is not conversational. You create a reusable avatar from a prompt, a profile or a photo for a one-time $0.95, then POST a script to /v1/avatar-1.0/talking-video and get an MP4 of up to 60 seconds. Per second the price is $0.184 (standard), $0.245 (plus) or $0.55 (max) without a product image, so a 30-second plus clip is 30 x $0.245 = $7.35.
If a product image is sent, the rates are $0.194, $0.258 and $0.58 per second.
A sensible split while you wait
- Render the fixed parts now: greetings, onboarding steps, product updates.
- Keep the interactive parts as text or a human until a live model reaches general availability.
- Do not build a launch on a preview with no price and no date.
- Reuse the same avatar handle so a later live product can match the face of your clips.
What to do
Ship the rendered clips this week and revisit when Tavus publishes access and price. Check the page again before you plan around it, because the availability line is the one most likely to change.
Sources
Related posts
More in Models
- Thai, Vietnamese, Indonesian voiceover: on MAI's list, not Sume's tags
MAI-Voice-2.1 lists th-TH, vi-VN and id-ID; Sume's 16 voice tags do not. What that means, a 1-cent audition, and the Unicode trap in Vietnamese.
- Translate text in an image but keep brand names: the edit prompt
A prompt pattern for translating the words in an image on Sume while leaving brand names, prices and codes alone, with an Ideogram 4.5 request.
- Turkish TTS API: MAI-Voice-2.1 tr-TR vs Sume tr, and the dotted I bug
Turkish is on MAI-Voice-2.1 (tr-TR) and in Sume's voice tags (tr). Upper-casing a script in code turns i into I; the failure and a one-cent test to catch it.
- /v1/videos audio_url references: Seedance 2 yes, Omni and Kling no
Only models whose supported_input_references lists audio_url take audio references on Sume: Seedance 2.x, Wan 3.0 and MiniMax H3 yes; Omni and Kling no.
Written by Sume