What Sume Avatar 1.0 Does and Does Not Do vs Live Avatars
A plain list of what Sume Avatar 1.0 renders (scripted 4-60 s clips) and what it does not do (real-time conversation), set beside Tavus Griffin's live model.
Sume Avatar 1.0 renders a scripted talking video from a reusable avatar. It does not hold a live conversation, listen to a viewer or react in real time. If you need a presenter who answers questions, this is not that product; if you need a clip you can review before anyone sees it, it is.
For context, Tavus's Griffin page, read on 2026-10-05, describes Griffin-Lite as a research preview of a full-duplex model that outputs 720p video in 320 ms chunks with 0.43 s audio-to-video latency on H100s. The page says it is not available to customers and is open only to select trusted testers, and it mentions no pricing or API.
What does it do today?
The table lists what Sume's own docs and rate card state, each read on 2026-10-05. Per-second prices are for the 720p output, without a product image.
| Topic | What Sume Avatar 1.0 does |
|---|---|
| Video length | Scripts or scene plans estimated at 4 to 60 seconds |
| Output | 720p; aspect ratios 1:1, 3:4, 9:16, 4:3, 16:9 (default 9:16) |
| Quality tiers | standard (fastest), plus (default), max (slowest); $0.184, $0.245, $0.55 per second |
| Creating an avatar | From a prompt, profile or photo; $0.95 once, handle reused after |
| Scenes | One avatar and one shared scene per video; silence beats need a duration |
| Captions | Optional inline captions, up to 60 seconds; a caption failure does not fail the video |
| Previews | First-frame preview stills before a full render |
| Face swap | Beta, separate route, source video about 4 to 15 seconds |
| Talking photo | Separate Fabric route: still plus your audio, 1 to 300 seconds |
| Script language | English scripts for Avatar Video; other languages go through Fabric with your own audio |
What does it not do?
It does not run a real-time or live conversation, and Sume makes no such claim. It does not listen to a viewer, take turns, or react to what someone says mid-clip. It does not render beyond 60 seconds in one job, so a longer script has to be split into several jobs, per Generate avatar video.
It also does not give you a measured duration in advance: the 4 to 60 second check uses Sume's estimate from the script, at 2.8 words per second rounded up per scene.
When should you pick which?
Choose a scripted render when the message is fixed: onboarding, ads, reminders, notices. Review a first-frame preview, render once and send it to everyone.
Choose a live model only if a viewer must be able to interrupt. At the time of reading, Griffin-Lite is a research preview, so there is nothing to buy yet. Check Tavus's page for changes, and keep your scripted clips as the dependable path. Disclosure matters either way: Tavus reports 48 percent of 54 study participants believed they were talking to a real person, so label an avatar as one.
Sources
Related posts
More in Sume Avatar 1.0
- Can I use Tavus Griffin-Lite yet? Preview status and what to ship
Tavus Griffin-Lite is a research preview for select trusted testers, not open to customers. Here is what the page says, and what you can build on Sume today.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
Written by Sume