What Sume Avatar 1.0 Does and Does Not Do vs Live Avatars

A plain list of what Sume Avatar 1.0 renders (scripted 4-60 s clips) and what it does not do (real-time conversation), set beside Tavus Griffin's live model.

5 min readSume
All posts

Sume Avatar 1.0 renders a scripted talking video from a reusable avatar. It does not hold a live conversation, listen to a viewer or react in real time. If you need a presenter who answers questions, this is not that product; if you need a clip you can review before anyone sees it, it is.

For context, Tavus's Griffin page, read on 2026-10-05, describes Griffin-Lite as a research preview of a full-duplex model that outputs 720p video in 320 ms chunks with 0.43 s audio-to-video latency on H100s. The page says it is not available to customers and is open only to select trusted testers, and it mentions no pricing or API.

What does it do today?

The table lists what Sume's own docs and rate card state, each read on 2026-10-05. Per-second prices are for the 720p output, without a product image.

Sume Avatar 1.0 capabilities (docs and rate card read 2026-10-05)
TopicWhat Sume Avatar 1.0 does
Video lengthScripts or scene plans estimated at 4 to 60 seconds
Output720p; aspect ratios 1:1, 3:4, 9:16, 4:3, 16:9 (default 9:16)
Quality tiersstandard (fastest), plus (default), max (slowest); $0.184, $0.245, $0.55 per second
Creating an avatarFrom a prompt, profile or photo; $0.95 once, handle reused after
ScenesOne avatar and one shared scene per video; silence beats need a duration
CaptionsOptional inline captions, up to 60 seconds; a caption failure does not fail the video
PreviewsFirst-frame preview stills before a full render
Face swapBeta, separate route, source video about 4 to 15 seconds
Talking photoSeparate Fabric route: still plus your audio, 1 to 300 seconds
Script languageEnglish scripts for Avatar Video; other languages go through Fabric with your own audio

What does it not do?

It does not run a real-time or live conversation, and Sume makes no such claim. It does not listen to a viewer, take turns, or react to what someone says mid-clip. It does not render beyond 60 seconds in one job, so a longer script has to be split into several jobs, per Generate avatar video.

It also does not give you a measured duration in advance: the 4 to 60 second check uses Sume's estimate from the script, at 2.8 words per second rounded up per scene.

When should you pick which?

Choose a scripted render when the message is fixed: onboarding, ads, reminders, notices. Review a first-frame preview, render once and send it to everyone.

Choose a live model only if a viewer must be able to interrupt. At the time of reading, Griffin-Lite is a research preview, so there is nothing to buy yet. Check Tavus's page for changes, and keep your scripted clips as the dependable path. Disclosure matters either way: Tavus reports 48 percent of 54 study participants believed they were talking to a real person, so label an avatar as one.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume