Full-duplex or script-driven avatar: six questions before you pick

Tavus Griffin-Lite is a full-duplex video conversation model in research preview. Six questions that separate it from Sume Avatar 1.0 talking video.

5 min readSume
All posts

Tavus Griffin-Lite, reported on 2026-10-01, is a full-duplex video conversation model, in research preview for invited testers (per a third-party tracker, read 2026-10-05; I read no more detail than that). Sume Avatar 1.0 talking video is script-driven: you send a script or an ordered list of scenes, and a job returns a finished clip. Before you pick, answer six questions. They tell you whether you need a live conversation, or a rendered clip.

The six questions

Question 1 is the one that decides the rest. If a person has to interrupt the avatar, a script cannot give you that. If the avatar says the same lines to everyone, a script gives you a file you can review, caption and reuse.

Full-duplex conversation vs Sume Avatar 1.0 talking video, with only facts from the Sume docs (read 2026-10-05)
QuestionIf the answer is yesWhat Sume Avatar 1.0 does
1. Does the viewer talk back while the avatar speaks?You need a live, two-way model.No. The input is a script or video_inputs. The result is a clip.
2. Is the exact wording fixed before release?A script is enough.Yes. Provide script or video_inputs, not both.
3. How long is the piece?Over 60 seconds means several jobs.Sume accepts a target duration of 4 to 60 seconds.
4. How many people are on screen?More than one means a different design.One resolved avatar for each final video.
5. What frame do you need?Set it per channel.aspect_ratio: 1:1, 3:4, 9:16 (default), 4:3, 16:9. resolution: 720p at this time.
6. Can you review before paying for the render?Previews help.avatar-video-previews returns first-frame stills, then generate-video renders.

What a script gives you that a call does not

A script-driven clip is a durable artifact. The result includes a public media.sume.com video, and may include preview_image_url and scene_previews. You can poll the job, read its events, and run the same text again with another quality tier: standard, plus (the default) or max.

You can also add silence beats. A scene with voice.type: "silence" has a required duration and no speech, so the avatar can hold a pause where a listener would nod. That is a scripted pause, not a reaction to a real person.

What a preview-only model leaves open

For Griffin-Lite, the only fact I can state is the research-preview status. Price, latency, resolution, API access and safety terms are things to ask the vendor. Do not fill them in from memory.

If a launch date matters to your plan, build the script-driven version first. The avatar, the script and the captions you make now will still be valid when a live option opens.

Decision in one line

  • Need two-way conversation: wait for access, and ask the vendor the open questions.
  • Need a finished clip for a feed, an ad or a lesson: use Avatar 1.0 now.
  • Not sure: render one 15-second clip, read its preview stills first, and see whether a clip answers the need.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume